Analysing the tools of 11 coding harnesses

Every tool in the toolbox of eleven open-source coding harnesses, catalogued by reading the registries out of the source code and normalised to a shared vocabulary of broad domains and finer-grained jobs so they can be compared.

Capability overlap between every pair of harnesses: how much of the smaller one's toolbox the bigger one already covers. The ringed cells are the thirteen pairs where the answer is all of it.

How many tools each one ships

Registry size, counted from source. Shipped tools, not tools enabled by default.

How much do they overlap?

Each cell asks how much of the smaller harness's capabilities the bigger one already covers. Overlap divides by the smaller set, so 100 means total containment. Click a cell to see what the pair shares.

The same tools, over and over

A tool's domain is the area it works in; its operation is the verb. A job is the pair, and it is the resolution at which two tools are really the same tool. Most registries ship close to one tool per job — they are wide, not repetitive.

Tools against jobs

The last column counts tools that duplicate a job the same harness already covers.

Who rebuilt what

Every job built by two or more harnesses, and what each one calls it. Read a row across: that is one tool, written independently, under a different name each time.

Capability coverage

The 43 domains against the eleven harnesses. The tally on the right is how many ship it — the shape of the consensus.

The registries

Pick a harness to read its full tool list and where it was read from.

The full dataset behind every table on this page is data.json.

Method, and what to distrust

Every inventory was read from vendored source except two: Kimi Code, whose runtime kimi-cli was not in the corpus, so its toolset comes from the SDK's own agent-file documentation; and Claude Code, read from a live session's tool manifest. Counts are tools shipped, not tools enabled by default — Codex, Qwen Code and DeepSeek all gate large parts of their registry behind flags, modes or agent permissions, so a default-set version of this analysis would look different.

The 43 capability domains were defined by hand to be the shared vocabulary, which means they are built to make equivalences visible. Treat the shape of the distribution as the finding, not any single percentage. Jobs split each domain by operation, which is finer and fairer for counting, but the operation for each tool is also a judgement call.