Forge Daughter

Most good work begins as a question.

This site is not one of those. It is a domain I own and use for experiments, and most days I could not tell you what it is either. A few of the experiments turn out fun. The rest get quietly deleted.

The fun one right now is the bots: little faces for agents, free to make and free to use.

Here are some agent harness benchmarks:

Terminal-Bench 2.0, accuracy
  • Codex CLI 82.2
  • Factory Droid 77.3
  • Junie CLI 71.0
  • Gemini CLI 61.4
  • Warp 61.2
  • Claude Code 58.0
  • Goose 54.3
  • OpenHands 51.9
Best entry per agent on the archived board, last entry May 2026. Source
Terminal-Bench 4.0, accuracy against cost
0%
20%
40%
60%
Codex
Claude Code
mini-SWE-agent
$0 Cost per trial $20
Best effort level per agent and model, September 2026. Cost is API spend per trial. Hover a point for its model. Source
Terminal-Bench 4.0, time
Claude Code 97.2
mini-SWE-agent 33.3
0 Minutes per trial 100
Average minutes per trial, same runs as above. Source
Terminal-Bench 2.0, best score to date
43.4
84.7
OctNovDecJanFebMarAprMay
October 2025 to May 2026, by entry month. Source