Forge Daughter
Most good work begins as a question.
This site is not one of those. It is a domain I own and use for experiments, and most days I could not tell you what it is either. A few of the experiments turn out fun. The rest get quietly deleted.
The fun one right now is the bots: little faces for agents, free to make and free to use.
Here are some agent harness benchmarks:
Terminal-Bench 2.0, accuracy
- Codex CLI 82.2
- Factory Droid 77.3
- Junie CLI 71.0
- Gemini CLI 61.4
- Warp 61.2
- Claude Code 58.0
- Goose 54.3
- OpenHands 51.9
Terminal-Bench 4.0, accuracy against cost
0%
20%
40%
60%
Codex
Claude Code
mini-SWE-agent
$0 Cost per trial $20
Terminal-Bench 4.0, time
Claude Code 97.2
mini-SWE-agent 33.3
0 Minutes per trial 100
Terminal-Bench 2.0, best score to date
43.4
84.7
OctNovDecJanFebMarAprMay