BANTAM FACTORY / fight card
Retry budget.
Respect backoff, server retry hints and a hard remaining-time budget without overflow.
Matchup completed with follow-up runs. OpenCode · 2026-09-08. Original BANTAM FACTORY result retained. Task, starter, grader and local model hashes match. Run conditions ↓
BANTAM FACTORY'S LOCAL SETUP
NVIDIA RTX 4090 · 24 GB
Qwen 27B · Q4_K_P · 72,192 context tokens
What does this mean for my setup?
This run used Qwen 27B on NVIDIA RTX 4090 · 24 GB. Compare generation speed using the same model and quantization. GPU memory, CPU offload, context size and prompt caching also affect performance.
Generation speed is measured while the server produces tokens. The task clock also includes reading context, running tools and testing. A GPU speed difference does not translate directly into the same change in total task time.
Fresh prompt processing: 1,651 tok/s · 23/23 timed requests. Rates use total recorded tokens divided by total server time for each phase. Hardware is operator-confirmed; timings come from the saved server responses.
The recorded factory floor
Meet the contenders.
4 local configurations.
3 frontier references, shown separately.
Choose any contender. Open its work.
Bars show elapsed time against the longest recorded run, not percentage of work completed. Token counters advance only when a saved response supplies them.
Recorded system comparison
Retry budget
| System | Outcome | Time | Groups | Project accepted | Clean finish | Input | Output | Cached input | Fresh input | Scope |
|---|---|---|---|---|---|---|---|---|---|---|
| BANTAM FACTORY · localQwen 27B · same local model | PASS | 141.7s | 5/5 | Yes | Yes | 304,594 | 9,969 | 273,649 | 30,945 | Recorded totals |
| DeepSeek HarnessQwen 27B · same local model | PASS | 445.8s | 5/5 | Yes | Yes | 397,922 | 27,046 | 384,287 | 13,635 | Recorded totals |
| HermesQwen 27B · same local model | PASS | 562.0s | 5/5 | Yes | Yes | 510,541 | 33,214 | 481,155 | 29,386 | Measured subset |
| Codex · AstraGPT-6 Astra · native CLI | PASS | 137.3s | 5/5 | Yes | Yes | 90,451 | 3,684 | 67,072 | 23,379 | Recorded totals |
| Codex · SolGPT-5.6 Sol · native CLI | PASS | 231.8s | 5/5 | Yes | Yes | 127,218 | 6,251 | 96,768 | 30,450 | Recorded totals |
| Codex · TerraGPT-5.6 Terra · native CLI | PASS | 143.2s | 5/5 | Yes | Yes | 163,987 | 5,762 | 137,728 | 26,259 | Recorded totals |
| OpenCodeQwen 27B · same local model | PASS | 375.4s | 5/5 | Yes | Yes | 433,649 | 22,979 | 384,597 | 49,052 | Recorded totals |
Qualification historyProgress includes the misses.0 local editions +
Selected qualification editions, not every development experiment. Earlier included failures are retained. These are separate attempts, not repeated trials of an unchanged system. An unrecorded task is not a failure.
Method & evidenceThe work order and run conditions.Read the method +
The recorded comparison
1 included work order: REPAIR — Retry budget (repeat 1). 7 systems and 7/7 recorded attempts. Every arm receives the supplied materials for its work order; independent acceptance checks remain separate from worker tests.
Local configurations and frontier references are distinguished in each lane. A frontier result is not a same-model comparison. All included outcomes remain visible; missing results are not zero-cost runs and do not count as passes. Elapsed times and completion status are reported separately.
These are selected development work orders. The original and follow-up recording windows, model settings, tool policies and time limits are documented with each card.
Reading an outcome
PASS: project accepted and clean harness completion. OUTPUT_ONLY: the artifact passed but the run did not reach accepted completion. FAIL: a required check failed. Timeouts and unrecorded tasks remain visible.
Passing independent groups alone does not override a failed public suite or protected-file check. Accepted project and clean finish are separate facts. The independent groups check the supplied work order.
Clocks & counters
Replay aligns each run's recorded start to zero. Runs did not all start simultaneously. Counters use actual saved response times, not animated estimates. Untimed evidence is not assigned an invented timestamp.
Input includes cached input; fresh input excludes cache hits. These three numbers are not additive. Partial metering displays only the measured subset, with coverage disclosed. Unknown never means zero. Native reports and server windows are overlapping scopes, not extra tokens.
The recorded work
The measurement download contains the task identity, outcomes, timings and token receipts used by this page. Reviewed work records provide the associated actions, checks and delivered files. Qualification history contains selected editions, not every development experiment; earlier included failures are retained.
The page runs without telemetry. Source hashes bind the displayed measurements to their saved records.
Source summary SHA-25609439f243e8b760e1bf86c7161eb8e53197e182652a4879b6760176be44b5d05
Put your model to work.
Your model.
A better factory.
Recorded results, ready to explore.