Spots

1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.

Kaggle Benchmarking Challenge Submission This is a submission for the Kaggle Benchmarking Challenge Time zones look like arithmetic. Add some hours, maybe cross midnight, done.

They aren't. Twice a year a local time

They aren't. Twice a year a local time disappears or happens twice. The US and Europe change their clocks on different weekends, so for a few weeks "New York is five hours behind London" is wrong. Some places sit on :30 or :45 offsets. Some countries changed their rules in the last few years, so a model trained on older text can be confidently out of date. So I built Wall-Clock Traps: 59 scenarios, each asked two ways.

Clean: "Local date and time in Chicago, USA

Clean: "Local date and time in Chicago, USA: 2026-03-08 02:30. What is the local date and time in London, UK at that same moment?"

Messy: "Backup job on the Chicago server is

Messy: "Backup job on the Chicago server is set for 2:30 am local on Sunday March 8th. London team wants to watch it run. What time is that in London?"

Same facts, same answer. (The answer is that

Same facts, same answer. (The answer is that 2:30 am never happens in Chicago that night. The clocks jump from 2:00 to 3:00.)

That gives 118 items in nine groups: basic

That gives 118 items in nine groups: basic conversions, US/EU gap weeks, odd offsets (Nepal, Chatham Islands, Lord Howe Island), the date line, flights that land on a clock-change day, countries that changed their rules recently, times that never happen, times that happen twice, and controls that look like traps but aren't. Design choices that mattered:

The answer key is code, not me. Python's

The answer key is code, not me. Python's zoneinfo computes every answer from the IANA time zone database. I checked 25 by hand.

Exact-match grading. Every prompt asks for a last

Exact-match grading. Every prompt asks for a last line like ANSWER: 2026-03-16 14:00, or ANSWER: NONEXISTENT / ANSWER: AMBIGUOUS. No LLM judge.

Wrong answers get sorted. Off by exactly an

Wrong answers get sorted. Off by exactly an hour is a DST mistake. Off by a day is a date-line mistake. Matching a country's old rule is a stale-knowledge mistake. Flagging a valid time as impossible is a false alarm. Four messy prompts carry a bad hint on purpose, like a CFO who "always just adds five hours." Gemini 3.7 Flash (Kaggle's default model) Gemma 4 26B A4B (small open-weights model)

Same prompts for everyone, one attempt per item

Same prompts for everyone, one attempt per item. I ran the benchmark twice: once on the Kaggle leaderboard, and once in an analysis notebook that keeps every answer so I could see why models missed.

News

1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.

Kaggle Benchmarking Challenge Submission This is a submission for the Kaggle Benchmarking Challenge Time zones look like arithmetic.

@spots #dev
Source: Dev.to
See more like this