BANTAM FACTORY / fight card
Job planner.
Order dependent work deterministically. Propagate failure without losing the plan.
The recorded factory floor
Meet the contenders.
Same Codex models.
Native CLI and BANTAM FACTORY.
Choose a pair. Open its work.
Bars show elapsed time against the longest recorded run, not percentage of work completed. Token counters advance only when a saved response supplies them.
Recorded system comparison
Job planner
| System | Outcome | Time | Groups | Project accepted | Clean finish | Input | Output | Cached input | Fresh input | Scope |
|---|---|---|---|---|---|---|---|---|---|---|
| Codex · AstraGPT-6 Astra · native CLI | PASS | 159.9s | 5/5 | Yes | Yes | 92,896 | 4,622 | 76,800 | 16,096 | Recorded totals |
| BANTAM FACTORY · AstraGPT-6 Astra · wrapped CLI | PASS | 109.0s | 5/5 | Yes | Yes | 56,116 | 2,896 | 44,288 | 11,828 | Recorded totals |
Qualification historyProgress includes the misses.0 local editions +
Selected qualification editions, not every development experiment. Earlier included failures are retained. These are separate attempts, not repeated trials of an unchanged system. An unrecorded task is not a failure.
Method & evidenceThe work order and run conditions.Read the method +
The recorded comparison
1 included work order: REPAIR — Job planner (repeat 1). 2 systems and 2/2 recorded attempts. Every arm receives the supplied materials for its work order; independent acceptance checks remain separate from worker tests.
Each model has a native Codex lane and a BANTAM FACTORY lane. Select matching model names to compare harnesses. All included outcomes remain visible; missing results are not zero-cost runs and do not count as passes. Elapsed times and completion status are reported separately.
These are selected development work orders. The original and follow-up recording windows, model settings, tool policies and time limits are documented with each card.
Reading an outcome
PASS: project accepted and clean harness completion. OUTPUT_ONLY: the artifact passed but the run did not reach accepted completion. FAIL: a required check failed. Timeouts and unrecorded tasks remain visible.
Passing independent groups alone does not override a failed public suite or protected-file check. Accepted project and clean finish are separate facts. The independent groups check the supplied work order.
Clocks & counters
Replay aligns each run's recorded start to zero. Runs did not all start simultaneously. Counters use actual saved response times, not animated estimates. Untimed evidence is not assigned an invented timestamp.
Input includes cached input; fresh input excludes cache hits. These three numbers are not additive. Partial metering displays only the measured subset, with coverage disclosed. Unknown never means zero. Native reports and server windows are overlapping scopes, not extra tokens.
The recorded work
The measurement download contains the task identity, outcomes, timings and token receipts used by this page. Reviewed work records provide the associated actions, checks and delivered files. Qualification history contains selected editions, not every development experiment; earlier included failures are retained.
The page runs without telemetry. Source hashes bind the displayed measurements to their saved records.
Source summary SHA-25680a60814ea8caf689d9b5670dd251d0e06da689527eb77f86a1125dbd563fea3
Put your model to work.
Your model.
A better factory.
Recorded results, ready to explore.