Stats
The creator of Gantry collects statistics on his own Gantry jobs and presents them here, to give an idea about how Gantry performs in real-life situations.
2026-06-18 - 2026-07-30
19 codebases 282 runs archived
409 milestones 2.4k sprints
13k agent sessions
457 plan 2.1k build 108 fix
66.4M tokens out 34k edits
105k turns 180k commands
Execute-agent peak context
2,116 build sessions - exact per-session peaksThe main goal of Gantry is to break work down into chunks small enough for agents to execute without running up large context windows. The key metrics are about context windows, and we focus on the peak number: the largest size the context window reached at any time during the session.
Across 2,116 autonomous execute sessions, the peak context reached stays under the 150k design target in 76% of them - and 89.9% stay under 200k. Only 0.9% of sessions (18 of 2,116) peaked over 300k.
Each histogram is that harness's own execute sessions, bin width 25k, final bar 300k+; bar height is scaled to the busiest bin in that harness.
Supporting roles
the aside - non-execute agentsExecution is the most important type of agent run and the most context-constrained. The bottleneck we are targeting. Here are the other roles and their context peaks.
| Role | claude | n | codex | n | median peak |
|---|---|---|---|---|---|
| execute | 114,468 | 861 | 106,464 | 1,255 | |
| plan | 82,376 | 206 | 47,406 | 251 | |
| review | 62,831 | 871 | 74,498 | 1,385 | |
| fix | 58,987 | 26 | 57,826 | 82 | |
| replan | 54,470 | 185 | 40,577 | 170 | |
| milestones | 43,867 | 76 | 28,590 | 122 |
Agent runtime
active working time per run and agent runtime by roleHow long runs and agents take. The whole-run figure is active time: stopped time is excluded, and parked time is excluded where the journal records it. The harness columns are each harness's median; median, mean and p90 pool both.
| Role | claude | codex | median | mean | p90 | n |
|---|---|---|---|---|---|---|
| execute | 9m 4s | 6m 3s | 6m 57s | 9m 43s | 18m 56s | 2,113 |
| plan | 4m 51s | 2m 6s | 2m 47s | 3m 32s | 6m 5s | 457 |
| review | 1m 44s | 2m 4s | 1m 55s | 2m 30s | 4m 17s | 2,218 |
| fix | 3m 39s | 2m 3s | 2m 11s | 3m 13s | 6m 40s | 106 |
| replan | 2m 16s | 1m 13s | 1m 45s | 2m 24s | 4m 17s | 229 |
Decomposition
269 structured buildsGantry sequences work by modelling it as milestones and sprints, and dispatching agents to implement sprint by sprint, and dispatching agents to do planning on multiple levels. Here is how the milestone-sprint metaphor ends up slicing the work.
| Metric | min | median | mean | p90 | max |
|---|---|---|---|---|---|
| Milestones / multi-ms job n=144 | 1 | 3 | 2.8 | 4 | 8 |
| Sprints / job n=269 | 1 | 7 | 8.8 | 18 | 57 |
| Sprints / milestone n=144 | 1.3 | 4.5 | 4.6 | 6.0 | 12.0 |
269 structured builds, bucketed by milestone count. A flat build is one whose sprints are a single list, with no milestone layer above them.
269 structured builds, bucketed by total sprint count.