ThoughtML for AI agents
ThoughtML is designed for an age where an AI agent can emit reasoning structure at no cost, and a human (or another agent, or CI) audits it. This guide is about that workflow.
Why a language, and why this one
When an agent explains a decision in prose, the explanation is unstructured: you can read it, but you can’t check it mechanically. ThoughtML gives the agent a target format that is:
- Cheap to emit. It’s plain text with a small, regular grammar. An LLM can produce it reliably.
- Typed and explicit. Every claim has a kind, every link a direction and meaning, every belief a confidence and (optionally) a basis. Nothing important is implied.
- Auditable. Once it’s structure, the mirror can read it a second way and flag where the agent’s own structure betrays its stated confidence.
The point is not to have the agent compute the answer in ThoughtML. It’s to make the agent’s reasoning legible enough that its flaws can’t hide.
The author/auditor loop
agent reasons → emits .thml → mirror reads it back → conflicts surface → human/agent resolves
A concrete version:
- An agent decides to ship a change and writes a ThoughtML document: the claim
(
cache-is-safe), the evidence it weighed, its confidence, and the basis of each number. - CI runs
thoughtml --audit(and maybe--strict-provenance). - The conflict report catches that the agent held a claim at 0.9 that its own recorded counter-evidence defeats.
- A human looks at exactly that one disagreement — not the whole paragraph.
Self-correcting against machine-readable diagnostics
An agent authoring ThoughtML doesn’t have to get it right first try — it can loop
against the validator. thoughtml check --json emits each diagnostic with a stable
code, the offending
line, and a suggested help fix (e.g. unknown relation → did you mean
supports?). The loop is: emit → check --json → apply the suggested fixes →
repeat until clean. Add --lint to catch the supports-used-as-a-list smell that
silently inflates confidence.
For authoring from scratch, the whole language travels inside the tool: run
thoughtml guide --full for a single self-contained, source-derived brief — closed
vocabularies, the distinctions that matter, and how to read the mirror back — meant to
be pasted into a system prompt. (thoughtml guide alone prints a one-screen tour, and
thoughtml guide <topic> looks up a single section.) It’s the same
llms.txt the site serves and the
packages embed — one source, so the CLI, the site, and the dump can never disagree.
Practical tips for generating ThoughtML
- Declare foci with explicit
kinds. It makes the graph readable and lets the kind-mismatch lint catch category errors. - Use confidence ranges for genuine uncertainty (
0.45..0.70) rather than a false-precision point estimate. - Always set a provenance basis. This is
the single most valuable habit for an agent: a
0.9 assumedis honest in a way a bare0.9is not. Run CI with--strict-provenanceto enforce it. - Record the counter-evidence. The mirror can only catch a
confidence-vs-statusconflict if the opposing observation is in the document. An agent that writes down what argues against its own conclusion gets the most value. - Keep computed and authored numbers separate — which the language does for you: never put a derived value in an authored field.
In CI
A minimal gate for agent-authored documents:
# fail on malformed structure (warnings included) and missing provenance
thoughtml --strict --strict-provenance reasoning.thml > model.json
# inspect the conflict report
thoughtml --audit reasoning.thml | jq '.audit.conflicts'
A non-empty confidence-vs-status error is a signal worth a human’s attention:
the agent believed something its own structure defeats.
See assistant-memory.thml for an agent’s evolving memory
of a user, and ship-the-hotfix.thml for the audit in
action.