Stanford paper: language models hold two separate notions of "the current year", and prompting fixes only one
"Do Language Models Consistently Encode the Current Year?" (van Adrichem, Bhaskar, Yang, Potts, Huang; arXiv 2608.15507, COLM 2026) finds that models guess "now" to within about a year of their training cutoff, and that telling them the date updates the year they state (94.6% success) but almost never the year they implicitly reason from (1.7%). This is a mechanistic account of why models with a stated date still act as if it were their cutoff year.
Key facts
- 13 models: base models predict a current year close to their post-training cutoff, with an average error of about 10 months
- Across 351 target years, prompting shifted the declarative (stated) year 94.6% of the time but the associative (implicit) year only 1.7%
- Year-shifted SFT moved the associative year in only 1 of 8 models; weight editing worked per task but did not generalise to both representations
- Submitted 2026-08-16; accepted to COLM 2026
What happened
The authors separate two things a model can "know" about the date: the year it says when asked, and the year built into its associations. They show that different mechanisms encode these, and that the usual fix of putting the date in the system prompt only reaches the first.
Why it matters
It explains a failure this dataset exists to reduce. A model told "today is 2026-09-29" can still treat post-cutoff events
as impossible or fictional. Background and related papers (Chunky Post-Training, chatbots as news intermediaries) are in
docs/cutoff-blindness/research.md.
Changelog
- 2026-09-29: created
Sources (1)
id: 2026-08-16-do-lms-encode-current-year · updated 2026-09-29 · open in the interactive timeline