skip to content

Stats

The creator of Gantry collects statistics on his own Gantry jobs and presents them here, to give an idea about how Gantry performs in real-life situations.

2026-06-18 - 2026-07-30

19 codebases 282 runs archived

409 milestones 2.4k sprints

13k agent sessions

457 plan 2.1k build 108 fix

66.4M tokens out 34k edits

105k turns 180k commands

Execute-agent peak context

2,116 build sessions - exact per-session peaks

The main goal of Gantry is to break work down into chunks small enough for agents to execute without running up large context windows. The key metrics are about context windows, and we focus on the peak number: the largest size the context window reached at any time during the session.

Across 2,116 autonomous execute sessions, the peak context reached stays under the 150k design target in 76% of them - and 89.9% stay under 200k. Only 0.9% of sessions (18 of 2,116) peaked over 300k.

Under 150k target
76%
1,599 of 2,116 sessions
Under 200k
90%
only 213 sessions above
Median peak
109k
across all execute sessions
Over 300k
0.9%
18 sessions
claude execute 861 sessions
0
25
50
75
100
125
150
175
200
225
250
275
300+
150k 200k
72%< 150k
114kmedian
13%> 200k
2.1%> 300k
542kmax
codex execute 1,255 sessions
0
25
50
75
100
125
150
175
200
225
250
275
300+
150k 200k
78%< 150k
106kmedian
8%> 200k
0%> 300k
246kmax

Each histogram is that harness's own execute sessions, bin width 25k, final bar 300k+; bar height is scaled to the busiest bin in that harness.

Supporting roles

the aside - non-execute agents

Execution is the most important type of agent run and the most context-constrained. The bottleneck we are targeting. Here are the other roles and their context peaks.

Per-role median by harness
Role claude n codex n median peak
execute 114,468 861 106,464 1,255
plan 82,376 206 47,406 251
review 62,831 871 74,498 1,385
fix 58,987 26 57,826 82
replan 54,470 185 40,577 170
milestones 43,867 76 28,590 122
Median over all roles: 76,673. claude peaks higher than codex in every role except review.

Agent runtime

active working time per run and agent runtime by role

How long runs and agents take. The whole-run figure is active time: stopped time is excluded, and parked time is excluded where the journal records it. The harness columns are each harness's median; median, mean and p90 pool both.

Runtime per agent role
Role claude codex median mean p90 n
execute 9m 4s 6m 3s 6m 57s 9m 43s 18m 56s 2,113
plan 4m 51s 2m 6s 2m 47s 3m 32s 6m 5s 457
review 1m 44s 2m 4s 1m 55s 2m 30s 4m 17s 2,218
fix 3m 39s 2m 3s 2m 11s 3m 13s 6m 40s 106
replan 2m 16s 1m 13s 1m 45s 2m 24s 4m 17s 229
Sprint runtime median
10m 38s
2,308 sprints - p90 25m 55s
Sprint runtime p25-p75
7m 2s - 16m 26s
execute + test + review per sprint
Active working median
1h 39m
whole-run active time
Active working mean
3h 24m
average across measured runs
Active working p90
6h 47m
90th percentile active time
Active working sample
282
runs with active duration
Stopped human waits are excluded for every active-duration observation. Parked retry waits are excluded for 0 newer runs; 282 older runs have stop-corrected active time only because their journals do not record parked waits.

Decomposition

269 structured builds

Gantry sequences work by modelling it as milestones and sprints, and dispatching agents to implement sprint by sprint, and dispatching agents to do planning on multiple levels. Here is how the milestone-sprint metaphor ends up slicing the work.

Sprints / milestone
4.5
median - p25-p75 3.7-5.3
Milestones / job
3
median of the 144 jobs that split into milestones
Sprints / job
7
median across all 269 builds - max 57
How a plan breaks down
Metricminmedianmeanp90max
Milestones / multi-ms job n=144 1 3 2.8 4 8
Sprints / job n=269 1 7 8.8 18 57
Sprints / milestone n=144 1.3 4.5 4.6 6.0 12.0
Milestones per job

269 structured builds, bucketed by milestone count. A flat build is one whose sprints are a single list, with no milestone layer above them.

flat 119
1 ms 6
2 ms 60
3 ms 47
4 ms 19
5 ms 6
6+ ms 4
Sprints per job

269 structured builds, bucketed by total sprint count.

1–2 sp 35
3–5 sp 86
6–10 sp 63
11–15 sp 36
16–25 sp 30
26–40 sp 9
41+ sp 2