The extent to which language models ignore word order remains a subject of debate in recent literature. Some studies find that altering word order has surprisingly little impact on certain downstream tasks, while others demonstrate that models actively utilize positional encoding and internal mechanisms sensitive to the sequence of tokens.
The evidence we hold leans evenly split
official record 3x · fact-check 2x · hedged 1x · crowd & reference 1x
The retrieved papers present conflicting evidence regarding whether transformer language models ignore or rely on word order. Papers [1] and [6] suggest that word order shuffling has minimal impact on model performance in certain contexts, supporting the idea that models can largely bypass fine-grained word order. Conversely, papers [2] and [9] demonstrate specific model sensitivities to token order and show that incorporating positional encodings significantly improves performance, indicating that word order is tracked and used. Because strong evidence is present on both sides, the verdict is CONTESTED.