I’m using a PyTorch Transformer model with identical preprocessing, tokenization, and weights for both training and inference. During training, the model produces stable and correct outputs, but during inference the predictions become inconsistent or nonsensical.
I have already checked the following:
model.eval()is set and dropout is disabledSame tokenizer, vocab, masks, and padding logic
Identical device/dtype settings (fp16 during training, fp32 during inference)
No missing or unexpected keys when loading weights
What subtle issues could cause a Transformer to behave correctly during training but diverge during inference? Are there known PyTorch-specific pitfalls related to positional encodings, attention masks, AMP scaling, or fused operations that might explain this?