Visual Reasoning Benchmark

Clock Bench

ClockBench evaluates whether models can read analog clocks - a task that is trivial for humans, but current frontier models struggle with.

Clock Faces
36
Clocks
180
Questions
720
Top Human Accuracy
95%
Top Model Accuracy
77.2%

Leaderboard

RankModelAccuracyLab
Top Human Performance95%
Average Human Performance90.7%
1Claude Opus 5.5 Max77.2%Anthropic
2GPT-6 Astra Max75%OpenAI
3GPT-6.1 Sol Max74.4%OpenAI
4GPT-5.6 Sol Max68.3%OpenAI
5Gemini 3.8 Flash66.7%Google
6Claude Opus 5 Max65.5%Anthropic
7Claude Fable 5.1 Max59.4%Anthropic
8Gemini 3.7 Flash53.3%Google
9GPT-5.4 High50.6%OpenAI
10GPT-5.5 High46.1%OpenAI
11Muse Spark 1.342.2%Meta
12Qwen 3-VL 235B Instruct39.4%Alibaba
13Claude Fable 537.2%Anthropic
14Gemini 3.5 Flash34.4%Google
15Gemini 3.1 Pro32.2%Google
16Gemini 3 Pro31.1%Google
17Grok 4.521.7%xAI
18Grok 4.620%xAI
19Gemini 2.5 Pro18.9%Google
20GPT-5.2 High15.6%OpenAI
21Gemini Robotics ER 1.515%Google
22Claude Opus 4.715%Anthropic
23Qwen 3-VL 235B Thinking14.4%Alibaba
24o3 Pro14.4%OpenAI
25o3 High12.2%OpenAI
26Gemini 2.5 Flash11.1%Google
27GPT-5 High11.1%OpenAI
28GPT-5 Pro11.1%OpenAI
29Mistral Medium 3.110.6%Mistral
30Claude Opus 4.610%Anthropic
31GPT-5 Mini8.9%OpenAI
32Claude Opus 4.18.9%Anthropic
33Claude Sonnet 4.57.8%Anthropic
34Claude Sonnet 46.7%Anthropic
35GTP-4o6.7%OpenAI
36Qwen 2.5-VL 72B6.1%Alibaba
37GTP-5 Nano3.9%OpenAI
38Grok 4 Fast3.3%xAI
ClockBench AI Benchmark
Top Human Performance
95.0%
Average Human Performance
90.7%
Claude Opus 5.5 Max
77.2%
GPT-6 Astra Max
75.0%
GPT-6.1 Sol Max
74.4%
GPT-5.6 Sol Max
68.3%
Gemini 3.8 Flash
66.7%
Claude Opus 5 Max
65.5%
Claude Fable 5.1 Max
59.4%
Gemini 3.7 Flash
53.3%
GPT-5.4 High
50.6%
GPT-5.5 High
46.1%
Muse Spark 1.3
42.2%
Qwen 3-VL 235B Instruct
39.4%
Claude Fable 5
37.2%
Gemini 3.5 Flash
34.4%
Gemini 3.1 Pro
32.2%
Gemini 3 Pro
31.1%
Grok 4.5
21.7%
Grok 4.6
20.0%
Gemini 2.5 Pro
18.9%
GPT-5.2 High
15.6%
Gemini Robotics ER 1.5
15.0%
Claude Opus 4.7
15.0%
Qwen 3-VL 235B Thinking
14.4%
o3 Pro
14.4%
o3 High
12.2%
Gemini 2.5 Flash
11.1%
GPT-5 High
11.1%
GPT-5 Pro
11.1%
Mistral Medium 3.1
10.6%
Claude Opus 4.6
10.0%
GPT-5 Mini
8.9%
Claude Opus 4.1
8.9%
Claude Sonnet 4.5
7.8%
Claude Sonnet 4
6.7%
GTP-4o
6.7%
Qwen 2.5-VL 72B
6.1%
GTP-5 Nano
3.9%
Grok 4 Fast
3.3%

Results Summary

Despite frontier models showing strong reasoning skills, mathematical ability, and visual understanding on multiple benchmarks, they seem to be struggling at reading analog clocks for now.

One hypothesis might be that this task sets a high bar for doing reasoning within the visual space (as opposed to text space).

More research is likely needed to understand if these capabilities can be obtained by scaling existing paradigms, or a novel approach is required.

Dataset

Sample Clocks

Few examples of clocks that we used in the benchmark.

Sample clocks from ClockBench

Questions

  1. Reading Time
    Models are asked to determine whether a given clock shows a valid time. If valid, they should report the hours, minutes, seconds, date, month, and day of the week (based on what is present), in a structured JSON format.
  2. Adding or Subtracting Time
    Models are asked to add or subtract varying amounts of time.
  3. Rotating Hands
    Models are asked to rotate one of the hands (hour, minute, or second) by a specified angle, clockwise or counterclockwise.
  4. Shifting Time Zone
    Models are asked to assume they are in New York during summer and report the corresponding time in various locations worldwide.

Try Yourself

Interesting in trying out ClockBench?
A small public dataset and sample evaluation code is available to everyone.

Public Dataset
Alek Safar
LinkedInX.com

Please reach out to [email protected] with ideas, suggestions, questions or any other inquiries.