BANTAM FACTORY / fight card
Snapshot drift.
Bind a snapshot to real files. Detect what changed, and validate what did not.
Matchup completed with follow-up runs. OpenCode · 2026-09-08; Pi · 2026-09-08. Original BANTAM FACTORY result retained. Task, starter, grader and local model hashes match. Run conditions ↓
BANTAM FACTORY'S LOCAL SETUP
NVIDIA RTX 4090 · 24 GB
Qwen 27B · Q4_K_P · 72,192 context tokens
What does this mean for my setup?
This run used Qwen 27B on NVIDIA RTX 4090 · 24 GB. Compare generation speed using the same model and quantization. GPU memory, CPU offload, context size and prompt caching also affect performance.
Generation speed is measured while the server produces tokens. The task clock also includes reading context, running tools and testing. A GPU speed difference does not translate directly into the same change in total task time.
Fresh prompt processing: 1,523 tok/s · 29/29 timed requests. Rates use total recorded tokens divided by total server time for each phase. Hardware is operator-confirmed; timings come from the saved server responses.
The recorded factory floor
Meet the contenders.
5 local configurations.
3 frontier references, shown separately.
Choose any contender. Open its work.
Bars show elapsed time against the longest recorded run, not percentage of work completed. Token counters advance only when a saved response supplies them.
Recorded system comparison
Snapshot drift
| System | Outcome | Time | Groups | Project accepted | Clean finish | Input | Output | Cached input | Fresh input | Scope |
|---|---|---|---|---|---|---|---|---|---|---|
| HermesQwen 27B · same local model | PASS | 497.0s | 5/5 | Yes | Yes | 211,676 | 31,045 | 196,714 | 14,962 | Measured subset |
| Codex · AstraGPT-6 Astra · native CLI | PASS | 146.2s | 5/5 | Yes | Yes | 98,851 | 3,822 | 87,552 | 11,299 | Recorded totals |
| Codex · SolGPT-5.6 Sol · native CLI | FAIL | 513.4s | 5/5 | No | Yes | 116,959 | 5,663 | 87,296 | 29,663 | Recorded totals |
| Codex · TerraGPT-5.6 Terra · native CLI | PASS | 106.2s | 5/5 | Yes | Yes | 145,342 | 4,101 | 106,752 | 38,590 | Recorded totals |
| BANTAM FACTORY · localQwen 27B · same local model | PASS | 124.3s | 5/5 | Yes | Yes | 409,391 | 8,124 | 377,438 | 31,953 | Recorded totals |
| DeepSeek HarnessQwen 27B · same local model | PASS | 465.8s | 5/5 | Yes | Yes | 485,350 | 28,346 | 470,652 | 14,698 | Recorded totals |
| OpenCodeQwen 27B · same local model | PASS | 403.2s | 5/5 | Yes | Yes | 265,821 | 25,282 | 224,796 | 41,025 | Recorded totals |
| PiQwen 27B · same local model | OUTPUT_ONLY | 600.0s | 5/5 | Yes | No | 642,107 | 38,643 | 613,734 | 28,373 | Measured subset |
Qualification historyProgress includes the misses.0 local editions +
Selected qualification editions, not every development experiment. Earlier included failures are retained. These are separate attempts, not repeated trials of an unchanged system. An unrecorded task is not a failure.
Method & evidenceThe work order and run conditions.Read the method +
The recorded comparison
1 included work order: EXTEND — Snapshot drift (repeat 1). 8 systems and 8/8 recorded attempts. Every arm receives the supplied materials for its work order; independent acceptance checks remain separate from worker tests.
Local configurations and frontier references are distinguished in each lane. A frontier result is not a same-model comparison. All included outcomes remain visible; missing results are not zero-cost runs and do not count as passes. Elapsed times and completion status are reported separately.
These are selected development work orders. The original and follow-up recording windows, model settings, tool policies and time limits are documented with each card.
Reading an outcome
PASS: project accepted and clean harness completion. OUTPUT_ONLY: the artifact passed but the run did not reach accepted completion. FAIL: a required check failed. Timeouts and unrecorded tasks remain visible.
Passing independent groups alone does not override a failed public suite or protected-file check. Accepted project and clean finish are separate facts. The independent groups check the supplied work order.
Clocks & counters
Replay aligns each run's recorded start to zero. Runs did not all start simultaneously. Counters use actual saved response times, not animated estimates. Untimed evidence is not assigned an invented timestamp.
Input includes cached input; fresh input excludes cache hits. These three numbers are not additive. Partial metering displays only the measured subset, with coverage disclosed. Unknown never means zero. Native reports and server windows are overlapping scopes, not extra tokens.
The recorded work
The measurement download contains the task identity, outcomes, timings and token receipts used by this page. Reviewed work records provide the associated actions, checks and delivered files. Qualification history contains selected editions, not every development experiment; earlier included failures are retained.
The page runs without telemetry. Source hashes bind the displayed measurements to their saved records.
Source summary SHA-2564676f4f7d1b2a628a4020ccb444d764c0413d2325d6ba103bc4e39a6b016dac0
Put your model to work.
Your model.
A better factory.
Recorded results, ready to explore.