ArgMap authoring tutorial
ArgMap authoring tutorial
Audience: a person who wants to read or write .argmap argument maps.
Format version: v0.2 (syntax frozen 2026-07-09; the v0.3 slash-pair
extension is noted where relevant). Semantics: ratified D36, amended by
D161 (the 2026-09-08 semantics campaign: the network reference as the
default solve, a count on every number, the kind key on lines, the set
rule for the tint, the what-if as revision).
Agent-facing companion: the argmap-author skill
(.claude/skills/argmap-author/SKILL.md) compresses this tutorial into
the working loop; agents load it via the Skill tool, humans can read it
as the cheat-sheet-plus. Since 2026-07-27 the skill is meant to be
sufficient on its own for authoring (it carries the idiom catalog,
the label/gloss and nesting discipline, the limitations and the lint
codes in compressed form), so the division of labour is: skill = the
pattern, this tutorial = the rationale, the history, the reading
chapter, and the worked example in Appendix A. A third tier,
examples/README.md, indexes the example corpus by idiom for when you
want to see a pattern in a whole file rather than as a fragment.
This tutorial distills the project’s design docs and the accumulated
authoring experience into one document. It never overrides them: on any
point of doubt, FORMAT_DESIGN.md (syntax), GRAMMAR_DRAFT.md (grammar),
SOLVER_SEMANTICS.md (semantics), and GLOSSARY.md (terminology) are
authoritative, and DECISIONS.md records why things are the way they are.
MATH.md is the readable account of the mathematics the semantics rests
on (what the numbers mean formally, what has been proved about them, and
what is still open), and is the right next stop after chapter 4.
AUTHORING_NOTES.md is the dated log this tutorial condenses; new
learnings continue to land there first.
Snippet convention: every .argmap block in this tutorial is either a
complete file that passes tools/argmap-lint.py as shown, marked
(complete, lintable), or an illustrative fragment whose first line is
# fragment - not standalone.
Contents:
- What ArgMap is
- Reading argmaps (self-contained; you can stop after this chapter)
- The format
- What the numbers mean
- From source text to map
- Labels and glosses
- Structural idioms
- Writing large maps
- Limitations
- Checking your map
- Appendix A: a complete worked example. Appendix B: cheat sheet.
1. What ArgMap is
ArgMap is a plain-text format plus an editor and viewer for making complex arguments explorable. Instead of reading a linear essay, a reader navigates the argument as a graph: the main claim and its support are visible at a glance, and every reasoning step can be unfolded to the depth the reader wants. The motivating use case is AI safety argumentation, where the arguments are long, branching, and full of objections that attack specific inference steps rather than conclusions. The flagship content is a comprehensive map of the book “If Anyone Builds It, Everyone Dies” (IABIED), deployed at p1graph.org.
An .argmap file describes a bipartite factor graph with two node kinds:
- Statements (written with the
@sigil) are variables: propositions that can be true or false, optionally annotated with the author’s credence that they hold. - Evidences (written with the
$sigil) are factors: reasoning steps that connect statements, optionally annotated with a reliability. An evidence line is one evidence written on one line of the file; readers see it as a reasoning link.
Roles such as premise, lemma, and conclusion are never declared; they are derived from the graph topology (a statement nothing points into is a premise, one nothing points out of is a conclusion). Attacks are not a separate primitive either: an objection is an ordinary evidence whose conclusion is a negated statement, and an attack on an inference (an undercut) is an evidence that references the attacked evidence itself. This uniformity is the core design idea: two node kinds and one reference mechanism express support, opposition, rebuttal, undercut, and refinement.
The text file is the single source of truth. The editor renders it as an outline and a graph, but everything those views show is derived from the text, and everything you author happens in the text.
The v0.2 syntax is frozen (DECISIONS.md D25 to D33). Anything this
tutorial shows is stable; future syntax changes arrive as versioned format
changes (the first is the v0.3 slash pair, gated by an explicit
argmap-version: 0.3 frontmatter declaration).
2. Reading argmaps
This chapter is for readers: people who explore existing maps in the viewer or query them from the command line. It does not assume or require anything from the authoring chapters.
2.1 The viewer
The editor/viewer at p1graph.org has three synchronized panes: the
text (the .argmap source), the outline (a collapsible tree of the same
content), and the graph. Reader and focus views present single nodes and
their neighborhoods in a more article-like form. The graph starts
collapsed: boxes with a fold control contain refinements, finer subgraphs
that replace a summary reasoning step when unfolded. Folding follows the
source structure, so what unfolds together is an authorial decision, not a
layout heuristic.
Conventions worth knowing when reading:
@nodes are claims;$nodes are reasoning steps between claims.- An evidence pointing at a claim supports it; an evidence pointing at a negated claim opposes it. An evidence that takes another evidence as an input attacks (or conditions on) that inference itself, not its conclusion.
- A number on a claim is the author’s asserted probability that it holds.
A number on a reasoning step is its reliability: roughly, how likely
the step is to actually carry when its inputs hold. A trailing
?marks a number as estimated rather than deliberately asserted; in the file only, since 2026-09-08: the popup’s numeral no longer carries the?, its sentence rides the hover text and the hollow caret on the gauge, and the popup shows instead a firmness beside every number, a count in coin flips on a claim and the step’s kind word on a reasoning step (2.2). - In the flagship map, numbers derive from the book authors’ own confidence language through a fixed rubric (DECISIONS.md D39), so disagreements the display surfaces are audits of the source’s coherence, not the map maker’s opinions.
2.2 Implied values and tension
The viewer can compute what all the authored numbers jointly imply. Under
“Show what the map implies” (in the Controls popover; on by default
since D40, though an explicitly persisted opt-out still wins), an
in-browser solver reads every authored number as a claim and finds the
distribution that honours all of them together, as far as they can be
honoured at once. Since D161 (2026-09-08) it starts from the map’s own
network of reasoning steps rather than from ignorance: each line is read
as a row of a probabilistic network, and the authored numbers are soft
targets on that network (before D161 the solve was the maximum-entropy
distribution honouring the numbers as constraints, each a spring rather
than a wall, D36). Each node
then shows an authored -> implied readout; the editor calls the
computed number the implied value, and the technical documents call the
same number the solved value.
Every number has a second axis. A number says how likely; beside it
the viewer shows how firmly it is held, in the currency of coin flips (a
number held at n flips gives way under pressure as far as a rate
estimated from n tosses would). The firmness is never a third thing to
author. For a claim it comes from the width its author left open: an
interval 0.7/0.1 (the claim somewhere between 0.7 and 0.9) is about 8
flips, a flat point 198 flips (about 200; the chip prints 198), which
is the cap every point gets. For a
reasoning step it comes from the step’s kind, read off the source’s own
words: a mechanism holds 64 flips, a record or a judgment 16, an analogy
or a hope 4, the source’s own deductive claim 1000, and a formal step
is hard. The popup shows the kind word on a step and the count on a
claim; the detail readout says the same in a sentence. Chapter 4 says
where the numbers come from (4.1 for steps, 4.7 for claims).
The tint means conflict. A readout colours only where the implied
value has left what the author wrote: below a step’s strength, outside
an interval, off a point, and by more than 0.01 (the solver’s own
resolution at a point, so a point met to 0.005 stays uncoloured).
Everything else is the solve filling what the author left open, and it
is shown without colour. On the flagship as authored that is the whole
picture: all 239 authored intervals are met inside their bounds, 112 of
them more than 0.05 above their lower bound (both counts re-measured
2026-09-24 on the shipped solve, solve_map.py at its default over
every pair statement; 224 and 113 on 2026-09-08, before the map’s later
content and the one draw), and none is tinted. A tint therefore says exactly one thing: the rest
of the map could not honour this number, and the size of the tint is
how far it fell short. Some statements carry a check credence (a
displayed comparison value that does not constrain the solve); the
badge comparing it to the implied value has the same meaning, the
distance from the implied value to the check’s interval. On the
flagship one sits past 0.10 as authored, and it is the headline:
@everyone-dies reads 0.603 against 0.75..0.95, 0.147 short, an audit
finding about the book; the next widest is its third premise @mis-ext,
0.795 against 0.85..0.95, a hair ahead of @fragile at 0.796 (measured
2026-09-28, after the two objections that deny @mis-ext were aimed at
it, D170; 0.617 on 2026-09-27, after the two inverses on the title claim
took out the reference fill that had held the headline at 0.708, D168). The headline’s check was re-read on 2026-09-25 from the book’s
unconditional sentences, with five other checks, each from the
sentence that prices its own quantity; the badge past 0.10 before
that, @evo-analogy’s, closed on 2026-09-08 when the chapter’s own
case for the analogy was mapped.
Moving a number yourself. In what-if mode a reader drags a claim’s number. The map re-solves with the reader’s number in place of the author’s on that claim, held as firmly as an author’s point (198 flips, the cap). On a claim nothing argues for, the map follows the reader forward and nothing colours. On a claim the author’s argument delivers, the reader’s number is a new claim beside the author’s case, and what bends shows what that case holds least firmly: whatever carries no number at all first, then, other things equal, in the order of firmness, hopes and analogies (4 flips), wide intervals, records and judgments (16), the author’s flat assertions (18 on the flagship), mechanisms (64), points (200), the source’s own deductive claims (1000), and a formal step never. The order is a tendency by count, not a rule by position: only numbers on the path between the reader’s claim and the case’s roots can give, so a mechanism on that path bends before a hope off it (4.6 shows one such case). Where the map cannot honour the reader’s number, the adjustment row says by how much: on the confusions map’s Socrates syllogism a “not mortal” at 0 reads “reached 32%”, the two premises giving to 0.67 each and the formal step holding (4.6, re-measured 2026-09-08 through the shipped solve; the row’s exact wording is of 2026-09-27, when “met at” became “reached”). Section 4.6 walks through both cases with the flagship’s numbers.
2.3 The headless readout
To query a map without a browser, use the CLI readout (from the repo):
cd experiments/solver-prototypes
python3 solve_map.py ../llm-extraction/iabied-comprehensive-en.argmap @shutdown '$link'
python3 solve_map.py ../llm-extraction/iabied-comprehensive-en.argmap --top 10
The first form prints, for each named node, the authored value and the
solved value. The second prints the ten largest gaps and tensions in the
whole map: the places where authored numbers and computed numbers disagree
most. --reference d36 --band adds the forced interval for a named
statement (how far the constraints actually pin it, as opposed to where the
solver settled inside the allowed range); it is a bench instrument of the
retired uniform reference, so it has to be asked for with that reference
and its numbers do not describe the shipped solve.
--condition if-built=0 --id @everyone-dies answers
a what-if by conditioning (the named statement taken as a fact about
the world, the author’s numbers updated by Bayes), the bench reading
that the viewer’s what-if mode does not use (4.6); --override is the
viewer’s reading, revision. Quote $id arguments so the shell does not
expand them. The solver needs python3 with numpy and scipy, and node on
the PATH.
That is everything a reader needs. To write maps, continue.
3. The format
An .argmap file is plain text, UTF-8, with optional YAML frontmatter,
comment lines, node lines, and footnote definitions. Indentation is
spaces only; a tab is a parse error.
3.1 Statement lines
@id [short label] p: gloss
@iddeclares a statement. IDs use[A-Za-z0-9_-], are case sensitive, and share one namespace with evidence IDs:@xand$xcannot coexist. Prefer mnemonic IDs (@risk-unbounded, not@s17).[short label]is optional. For statements the label is the claim, phrased as a proposition. If no label is given, the gloss serves as the display label.pis optional: the author’s credence that the statement holds, a probability literal in [0,1]. A trailing?(as in0.7?) marks the value as estimated or unelicited rather than deliberately asserted.- Everything after the
:is the gloss: one logical line of free text giving depth, sourcing, or qualifications.
A statement is a variable, so its label must be a proposition, something that can be true or false. “Anyone builds it” is a statement; “the question of whether anyone builds it” is not. Conditionality lives in evidences, never in statement labels. Normative propositions (“X should happen”, “doing Y is impermissible”) are legal and ordinary statements; what chapter 9’s fact/norm caution forbids is not norms but future-fact nodes that would feed back onto their own antecedents.
3.2 Evidence lines
$id [label] strength <conclusion-expr> | <premise-expr>: gloss
$iddeclares an evidence: a reasoning step asserting that its premises bear on its conclusion.[label]is optional and carries the headline warrant: why the premises support the conclusion, in a phrase (chapter 6).strengthis optional: a bare probability literal, the evidence’s reliability (chapter 4 explains precisely what it means).?works as on statements.- The
|is the given bar and reads “given”:$e @c | @ais evidence about@cgiven@a. The conclusion side is a full expression, not just a single reference. - The
| <premise-expr>part may be omitted entirely. A premise-less evidence is an unconditional constraint factor: it asserts its conclusion expression with the given reliability, unconditionally. Example from the spec:$rivals ~@hyp-fluke OR ~@hyp-filter: rival explanations can't both hold.
3.3 Expressions
Premise and conclusion sides use the same grammar:
- References:
@idfor a statement,$idfor an evidence (see 3.6), each optionally negated with~(~@id). Negating an evidence reference (~$id) is forbidden: it parses, but the validator rejects it (error E3). To challenge an inference, write an undercut (3.6). ANDjoins linked premises: the step needs all of them.ORjoins convergent premises: any one suffices.- Mixing AND and OR requires parentheses:
(@a AND @b) OR @c. Unparenthesized mixing is a parse error; there is no silent precedence. - Unicode
∧ ∨ ¬are accepted as input aliases; the canonical form is ASCII.&is not a connective (the|character is taken by the given bar, soORcannot be written|either).
Note the graph-level route to convergence: two separate evidence lines with the same conclusion are independent factors, which is usually the right way to say “two independent reasons” (see 4.4 for when it is not, and for what happens to two reasons once their conclusion is known).
3.4 Refinement (nesting)
Indentation always means “belongs to the line above”. What belongs
means is read off the parent line’s sigil: under @ and $ it is
refinement (this section); under :: it is membership in a
declared group (3.10). Everything below is the @/$ case.
An indented block under an evidence is a refinement: a finer-grained subgraph that models the same reasoning step at higher resolution. When a reader unfolds the evidence, the block replaces it; folded, the outer line serves as the coarse summary. One consistent indentation increase per level (two spaces recommended).
# fragment - not standalone
$syllogism @socrates-mortal | @socrates-human AND @humans-mortal: surface form
@intermediate [Socrates inherits mortality property] 0.99:
$inherit-1 @intermediate | @socrates-human AND @humans-mortal-property:
$inherit-2 @socrates-mortal | @intermediate:
A statement may also carry an indented block; that refines the implicit factor asserting the statement’s own marginal, and the statement itself remains.
Refinement changes what the folded line reads: a folded evidence shows the composition of the block under it, so the steps you write below a coarse line decide the number the reader sees on it. 7.5 says how to pick the coarse strength so the two agree.
IDs are document-global: a node declared inside a refinement can be referenced from anywhere, and forward references (using an ID before its declaration) are legal. Where you nest is a real authoring decision, not formatting: nesting determines what folds away together in every view (DECISIONS.md D22).
3.5 Glosses and continuation lines
A gloss is one logical line, but it can be hard-wrapped: any
deeper-indented line that does not begin with @, $, or # folds into
the gloss of the nearest preceding node line, joined with a space. This
is also how longer narrative passages attach to a node without costing
graph structure:
# fragment - not standalone
@human-precedent [Human intelligence transformed the planet] 0.9?: Nobels to humans, none to chimps
a parable: a council of beast-"gods" laugh at the Ape-god's newest
creature, frail and clawless. The Ape-god says only, quietly, "and yet."
The hazard: if you forget a sigil on a node line, the line silently becomes gloss text of the node above. The lint warns when a prose line looks like a node declaration (W3); take that warning seriously.
The inverse hazard has no warning, and cannot get one. A continuation
line that begins with @, $, #, :: or > is read as that
construct, not as prose, because line dispatch is a first-character
switch, and by the time anything could complain, the parser has built a
node and has no idea prose was intended. So never start a continuation
line with a sigil character: begin with a word, or rephrase. The exposure
is small because quote lines (3.12) cannot wrap and the corpus barely
uses continuation lines at all, but when it bites there is no
diagnostic: you find it by reading the rendered gloss.
Two lexical restrictions: a gloss cannot contain a # preceded by
whitespace (that always starts a trailing comment), and a label cannot
contain square brackets.
3.6 Evidences as premises: conditioning and undercuts
An evidence named in another evidence’s premise expression denotes that evidence’s activation (“this inference is in force”), not its conclusion. There are two uses.
The positive use is conditioning on an inference: $policy @act | $link
makes a conclusion depend on an implication holding, rather than on a
fact. This is rare and legal; the lint flags it (W1) because the same
shape is usually a polarity mistake, so when you do it deliberately, say
so in a comment.
The common use is the undercut. To attack the inference $E C | P
(rather than its conclusion), write:
# fragment - not standalone
$E-undercut ~C | <grounds> AND $E: why the inference fails
The conclusion negates $E’s conclusion; the premises conjoin the
grounds with $E itself. Conditioning on $E is exactly what makes this
an undercut rather than a rebuttal: if $E is itself disabled or
undercut elsewhere, the undercut lapses with it. Dropping the AND $E
turns it into a plain rebuttal, which fires regardless. Undercuts of
undercuts (reinstatement) are the same schema applied again.
Coming from probabilistic conditional logic
If you know conditionals of the form (psi | phi)[d] from the
maximum-entropy literature (Kern-Isberner, Paris, the SPIRIT and MEcore
systems), five translation rules keep the solve honest; each was measured
on examples/toys/pcl-penguin.argmap (2026-08-31, and again under the
shipped solve 2026-09-24; in the app, the Semantics audit map
audit-prior-art, group P1).
(psi | phi)[d]withdat or above 0.5 is the line$e d psi | phi.- With
dbelow 0.5 it is the opposed line at1 - d:$e (1-d) ~psi | phi. A line’s number is the inference’s reliability, so0.01 @fly | @pengis a nearly worthless reason for flying, and the solve concludes from “birds fly” that penguins fly (0.95 under the what-if). - A subclass exception is an undercut of the general rule on the
subclass plus a rebuttal.
0.9 @fly | @birdbeside0.99 ~@fly | @pengis a contradiction here, because “birds fly” in force is a law over penguins too; with the penguin pinned the map is infeasible. Write0.99 ~@fly | @peng AND $birds-fly(the undercut, which alone leaves flying at even odds) and0.99 ~@fly | @peng(the rebuttal, which then reads the textbook 0.01). - An independence statement, a conditional at the base rate such as
(allergic | treated)[0.10]beside(allergic)[0.10], has no line form. Leave it out: a support line at the base rate pulls its antecedent down, and the lint flags the attempt (W25). If the pull it was meant to block is real, pin the statement it protects. - A fact
(a)[d]is the bare point@a d; whether it should be a pair is the two-floor discipline of 4.7 (W23, W25).
The solver never reads your descriptions, so the same three lines with different glosses solve to the same numbers. What a description can do is tell you which shape to write. Three readings of “birds fly, 0.9” and the shape each licenses:
- The law. “All birds fly; I am 90 percent sure.” Write
0.9 @fly | @birdand, for every exception you know, the undercut plus rebuttal of rule 3. The rule stays a law over the whole class. - The frequency. “90 percent of birds fly.” Either scope the rule to
the class it is true of,
0.909 @fly | @bird AND ~@peng(the number raised so the mixture over the class comes back to 0.9) with@peng’s base rate written down, or keep the unscoped line and read the exception off the solved map by conditioning (the CLI’s--condition), never by pinning the exception: a pinned penguin beside an unscoped “birds fly” is a contradiction in this reading as much as in the textbook’s. - The case. “This bird probably flies.” A point or pair on
@flyitself, or a pinned premise; no rule about birds is being asserted.
3.7 Comments and comment-layer conventions
A line beginning with # is a comment; a whitespace-preceded # starts
a trailing comment. Comments are preserved by the parser and serializer.
Section headings in large maps are full-line comments by convention. Use
one when the heading is only for a human reading the source. When you want
the heading to be checked and drawn, a named box around those nodes,
use a declared group instead (3.11): no tool can see a comment.
Two trailing-comment conventions carry meaning to the solver tooling without being syntax:
# check: pon a statement line records the author’s all-things-considered credence for display against the computed value. It never constrains the solve. Chapter 4 explains when to use it instead of an authored marginal. A check value may carry the?marker like any other value (# check: 0.9?), and in a source-faithful map it should: there the check is the source’s own stated register for that conclusion (the D39 practice), not the extractor’s belief. A check may also be an interval,# check: 0.85..0.95: the range the register licenses, at the same widths the residual pairs use. The badge a reader sees is then the distance from the computed value to that range, zero inside it; a bare point is the zero-width interval. A token that is neither (reversed bounds, a bound past 1, a single dot) is the lint’s W26.# gate: q($id) >= 0.10 => @conclusionrecords a threshold audit: after a solve, if the left side clears, the named conclusion is expected to hold, and the display reports agreement or disagreement.# kind: <word>on an evidence line records what sort of step the line is, read off the source’s own words:formal,deductive,mechanism,empirical,testimony,analogyorhope. It sets how firmly the solve holds the line under a reader’s what-if (its count, in coin flips), never its strength; 4.1 says when to write which. An empirical line may add its stated sample size (# kind: empirical n=200), and the key may share a comment with other keys, separated by;(# check: 0.9?; kind: mechanism). The key is content: it is identical in every language of a translated set (TRANSLATION_NOTES L16, checked bytranslation-parity.py). On a refined coarse line the key is for the reader; the solve reads the leaves’ keys, because the refinement replaces the coarse line (3.4, 7.5).
All three are conventions, not grammar; tools other than the solver readouts will treat them as ordinary comments.
3.8 Citations
Attach citations as Markdown-style footnotes: [^ref] in a gloss, and a
definition line anywhere at top level:
# fragment - not standalone
@p1 [Capabilities advance rapidly] 0.9: doubling times keep shrinking [^epoch2025]
[^epoch2025]: Epoch AI, "Trends in Machine Learning," 2025.
Footnote text is free text: author, title, venue, year, plain URL. Viewers autolink URLs; there is no inline link syntax. The lint checks that every used footnote is defined and every defined footnote is used (W4).
3.9 Frontmatter and file layout
Optional YAML frontmatter between --- fences carries metadata: title,
author, date, description, source, scope (what part of the
source the map claims to cover, checklist item 6), and argmap-version
(declare 0.3 if the file uses slash pairs or quote lines). Unknown keys
are preserved, which makes frontmatter the extension point for
provenance notes.
One key is tooling-visible: focus: [id, id] (D57) declares the map’s
focus nodes, the statements influence readouts measure deltas on. Omit
it and the tooling derives them from topology (statements concluded by
top-level evidence and premised by none). Declare it only when topology
misreads your intent: the known case is a goal guard, a world-layer
conjunct like ... AND @shutdown on strategy-advice lines, which makes
the goal premise-referenced without arguing from it. The list is
complete, not additive.
One key is a display hint: fold-links: on (D162) asks the viewer to
open your map with its “Fold links” mode on, so a folded branch shows
which other branches it touches. Declare it when your map’s branches
cross each other and a reader needs to see that at a glance; leave it
out otherwise, since the mode is off by default and it does reshape the
folded layout. A reader’s own checkbox in the graph controls still wins
for their session.
A second display hint belongs to multi-voice maps: declines: (D164)
lists, per speaker key, the claims that speaker refused on the record
to put a number on, as a: everyone-dies [^t012346], ids without the
@. In that speaker’s view the claim then shows a blank gauge with
their words where the view’s arithmetic would stand. Use it only for a
spoken refusal, never for a claim the speaker merely left unpriced:
every entry needs a verbatim > quote line by that speaker on that
statement, and the optional [^locator] picks which one the reader
sees. The solve ignores the hint, so everything downstream of the claim
keeps its value in that view.
Top-level order is free; the graph defines the structure. Convention: put the document’s headline claim first, then work down its support.
3.10 v0.3: two-sided pairs (brief)
Since D52/D53 a file declaring argmap-version: 0.3 may write two-sided
values: on an evidence, $e 0.9/0.2 @c | @a adds an opposed floor toward
the negated conclusion in the same slab; on a statement, @s 0.8/0.1
bounds P(s) to [0.8, 0.9] instead of pinning a point. No whitespace
around the slash; ? binds per member; an omitted second member is 0 and
means exactly the v0.2 reading. New maps can ignore pairs until they need
to express “this consideration cuts both ways” or an interval-shaped
residual; details in FORMAT_DESIGN §3.1/§3.2 and SOLVER_SEMANTICS §1.9.
3.11 v0.3: declared groups (::)
A third sigil declares a group: a named box drawn around nodes.
::timelines [Capability timelines]: when transformative AI arrives
@agi-soon [Transformative AI within a decade] 0.4:
@compute-grows 0.9: frontier training compute keeps growing
$scaling 0.7 @agi-soon | @compute-grows: the trend argument
Its indented block is membership, not refinement (the one place the
indentation rule is keyed on the parent’s sigil, 3.4). Otherwise the head
reads exactly like a node line: required id, optional [label], optional
gloss, and the same continuation-line folding.
Three properties define it:
- No credence. There is no probability slot on a
::line, now or later. A number there is an error. - Not referenceable.
::idin any expression is a parse error. A group takes no part in inference: you cannot argue from it or against it. - Transparent. Deleting every
::line changes nothing about the map: same graph, same roles, same solve. A group is display only.
Use one to say “these nodes are one topic”. Before ::, the only way to
say that was a # ==== banner comment, which no tool could see, check,
or draw.
- A kind box is headed by its thesis (Felix’s ruling, 2026-09-09).
A group that boxes objections of one kind takes as its label the
shared thesis of its members, in the objector’s voice and general
enough that every member is an instance of it:
::halt-enforcement [A halt cannot be made to stick]. A reader arrives with a proposition in their head and matches by content; a neutral question head (Can the halt be made to stick?) makes them translate first. Check every member against the head before committing; where one is not an instance, adjust the head, never the member. The gloss opens on the member list, because the folded card’s popup shows it and it is the second discovery channel. The scope, decided the same day: the rule binds a KIND card nested under an objection shelf; a document-level shelf keeps its question or topic head (Is the worry real and near?,Objections to the halt), which reads well as it is and has no single thesis to state.
Groups and blocks. A block is derived: a connected component of the graph, a set of nodes that reach each other. A group is authored. They usually coincide, and the validator checks the relationship: a group equal to one block, or spanning several whole blocks (“two topics under one heading”), is silent. Two shapes warn, because your claim and the graph disagree:
- a group covering only part of a connected block (W12): edges cross its boundary, so the open box distorts the layout and a fold can only summarize the crossings;
- one block split across two groups (W13): usually an accidental cross-topic premise silently merged two topics while your headings still assert they are separate. This is the mistake worth catching.
Only document-level groups are checked. A group nested inside another group, or inside a refinement, is subdividing its parent, not claiming a block.
Folding. Every group folds to a card. A closed group (no edge crosses its boundary) folds with nothing to redraw; a non-closed group folds too, and its boundary edges become dashed summary edges bundled by direction and kind, so a battery of twelve objections folds to one card with one bundled attack edge. The card shows the label, a tally of what its hidden edges do to the surrounding claim, and never a credence: the summary vocabulary is the guard here, since the card only reports where hidden traffic goes and takes no part in the argument itself. A group nested inside a refinement starts folded, so opening a box shows its groups as cards first; a document-level group starts open.
When to group. The workhorse case, measured on the flagship map, is the objection battery inside a box: several answered attacks on the box’s claim, each objection paired with its response. Grouping them turns a wall of interleaved nodes into one labeled card next to the claim, and the claim’s supports become findable again. Guidance from that campaign: keep a family’s lines contiguous so the group is a pure wrapper; put a ground inside the group only if nothing outside the family consumes it, and leave shared grounds outside as boundary premises; give the group one dominant story toward its claim (a box that both attacks and supports its claim in equal measure makes an illegible card); label it as a short reader question (“Is the worry real and near?”) rather than a category noun. Leave a box flat when its objections interleave with shared-ground supports, and treat a group created purely to fix layout as provisional: re-judge it in the reader’s default view whenever the viewer changes.
Groups are allowed anywhere: any depth, inside each other, inside refinements. Membership does not suppress the isolated-statement note: a context shelf of standalone facts still reports each one as isolated.
3.12 v0.3: source quote lines (>)
A line beginning > under a node carries a verbatim span from your
source, plus the footnote locator it came from:
# fragment - not standalone
@no-honor [Honor is a contingent evolved hack an AI won't carry] 0.9?: honor is
an evolutionarily contingent shortcut, not a convergent feature of minds
> a specific weird hack that humanity stumbled into [^supp-ch5]
> quite skeptical that gradient descent will happen to stumble across the
(the last line is shown truncated only for the page width, see “no wrapping” below.)
The gloss goes back to being a claim a reader can parse cold; the quotes
sit under it as its evidence. Before >, a quote could only live inside
the gloss, where no tool could see it. That meant no display affordance,
and, in a translated map, nothing stopping a paraphrase from being
presented as verbatim.
Rules, all short:
- The text is verbatim. Never paraphrase it, never silently repair it. If you need to trim, trim at the ends.
- Always give a locator.
[^ref]at the end of the line, defined at top level like any footnote (3.8). A quote without provenance is almost always an authoring slip, and the lint says so (W15). Locators are as coarse or fine as your source allows: a chapter ([^epub-ch7]), a supplement page ([^supp-ch5-promises]), a transcript timestamp ([^t001734]). - Only a locator at the very end of the line counts. Everything else
on the line is verbatim text, including a
[^…]in the middle of it (W16 flags that as a probable stray or doubled ref). - No trailing
#comment, the one line kind that has none. Source text cannot be reworded to dodge the comment splitter, so a real ` # ` in a quote would be silently truncated; instead the whole line is verbatim and the lint warns if it spots ` # ` inside one (W17). Put per-quote notes on an annotation comment instead (below). - No wrapping. A quote is exactly one line, however long; the editor soft-wraps it.
- Placement is positional. A quote attaches to the node above it and
must be indented deeper. It has to sit in that node’s annotation
block, the span before the node’s first child. A
>at top level, or after a child node, is an error (E9), not a re-attachment to some outer node. Order quotes after the gloss prose; interleaving parses, but the lint prefers the canonical order (W18) and the serializer rewrites to it anyway. - Declare
argmap-version: 0.3in a file that uses>(W19).
Which quotes become > lines: the three-way test. Ask: is this the
node’s own wording, or support for it?
- Supporting quote: a fragment stacked next to the claim as
evidence for it. Lift it to a
>line. This is most of them. - Load-bearing inline fragment: a verbatim phrase that is a
grammatical constituent of the gloss sentence (
Kelvin's "infinitely beyond…" fell to DNA). Leave it in the gloss, in plain quotation marks: it is the node’s own phrasing, borrowing the source’s words. When the provenance is worth keeping, add an echo, a>line carrying the full verbatim sentence and its locator, while the gloss keeps its fragment. - The quote is the claim: the gloss is nothing but the quote.
Degenerate case of 2: write the gloss in plain marks and echo the
verbatim on a
>line.
The echo pattern also keeps translations honest: a translated gloss
renders the fragment as ordinary quoted prose (claiming nothing about
verbatimness), while the > line stays in the source language.
Quotes are never translated. In a multilingual map set the whole >
line (sigil, indent, text, locator) is byte-identical across all
language versions, and tools/translation-parity.py enforces that. A
translated “verbatim” quote is false on its face and destroys the tie
back to the source.
Annotation comments (#[…]). Per-quote side data goes on a full-line
comment of the form #[key: …] (no space between # and [) on the
line above the quote, at the same indent:
# fragment - not standalone
#[de: schwer, der Schlussfolgerung zu entgehen]
> hard to avoid the conclusion [^supp-ch5]
To the parser this is an ordinary comment. Two things make the form
worth using rather than a plain #: the parity tool treats #[…] lines
as free per file (every other comment must match byte-for-byte across
translations), and it is the reserved surface for real attributes in a
later format version, so today’s convention promotes without a rewrite.
Its current tenant is the parked translation of a quote, waiting for a
real translation field.
4. What the numbers mean
The numbers are the part of the format most worth getting right and the part where intuition most often misleads. The ratified semantics (DECISIONS.md D36, full treatment in SOLVER_SEMANTICS.md) reduce to a small set of rules an author can hold in their head.
This chapter gives those rules operationally: what to write, and why it
behaves as it does. If you want the model underneath them (what the map
compiles to, why the solve is a maximum-entropy problem, and which of
these rules are theorems rather than conventions), that is MATH.md,
published as The mathematics behind ArgMap. Nothing here depends on
reading it.
4.1 Evidence strength
Elicit an evidence’s strength by asking: assume the premises hold; how likely is the conclusion? That prompt is the whole elicitation procedure. Formally the number is the unconditional in-force rate of the rule (how often this kind of inference actually carries), a property of the rule itself, independent of whether its premises happen to be true. The two readings coincide numerically by construction, so you can elicit with the conditional prompt and reason with either picture.
Practical consequences:
- The strength isolates the inference. Whether the premises are true is carried by the premises’ own numbers, elsewhere in the map. Do not discount a strength because you doubt the premises.
- Contraposition is not a rewrite.
$e 0.8 @c | @aand$e2 0.8 ~@a | ~@care different claims; the given bar is directional. When extracting or translating, preserve the direction the source actually asserts. - A strength of 1 is legitimate for deductive steps: the line becomes a pure constraint (“a proved implication has no reliability coordinate”). A strength of 0 is almost never what you want; the unstrengthed line is the exact “structure only” form (the lint suggests this, W11).
- An evidence with no strength at all contributes structure to the display but nothing to the solve. This is a deliberate, useful state: sketch the shape of the argument first, commit numbers later.
- A near-certain premise costs nothing. Under the retired uniform
reference a line’s strength was held on both sides of its premise, so
a line conditioned on a near-tautology (for example the OR of four of
five partition members) showed a large spurious tension on the nearly
empty side, and this item warned against it. The shipped solve holds
a line’s rate once, on its own coin: a 0.8 line on a premise pinned
0.98 reads its strength exactly, where the retired reference read the
empty side 0.079 off (a two-line toy,
solve_map.pyat its default and at--reference d36, 2026-09-24). Condition on the premise the source states.
The second axis: the line’s kind. A strength says how likely the
step is to carry. Since D161 (2026-09-08) every line also has a
count: how firmly the solve holds that strength when something
else in the map pushes against it, in coin flips (a strength held at n
flips gives way as far as a rate estimated from n tosses would). The
count is never elicited as a number. It comes from the line’s kind,
what sort of step it is, which a reader can classify from the source’s
own words where nobody could read a count off a book. Write it as the
comment key # kind: <word> (3.7). The rubric, with the cue words and
the count each kind holds, as fixed on 2026-09-08 (deductive raised
from 200 to 1000 flips the same day, after the campaign’s skeptic pass;
DECISIONS.md D161 carries the current table, this one is dated):
| kind | the step | cues in the source | flips |
|---|---|---|---|
formal |
logic, definition, arithmetic, a machine-checked derivation, a universal instantiation; checkable independently of the author | a proof, a calculation, “all F are G and x is F” | hard (or simply write the step at 1) |
deductive |
the source’s own claim that the conclusion follows | “by definition”, “necessarily”, “it follows”, “is the same as” | 1000 |
mechanism |
a causal or structural reason that would operate whenever the premises hold | “because”, “the process”, “would tend to”, “any such system” | 64 |
empirical |
a frequency or a record | “historically”, “in every case so far”, a named count of incidents, a study | 16, or a larger stated sample size |
testimony |
the authors’ or experts’ stated judgment, offered as such | “we think”, “experts”, “most researchers”, “our judgment” | 16 |
analogy |
an inference carried by a likeness | “like”, “as with”, “analogous”, “the way evolution” | 4 |
hope |
the source’s own labelled hopes and speculative objections | “the hope:”, “perhaps”, “one might hope”, “it is conceivable” | 4 |
| no key | 16 |
Four rules ride on the table:
- Formal against deductive. A formal step is written at strength 1,
or at its strength with
# kind: formal; either way the solve holds it hard, a constraint no other content bends. A source’s own deductive claim is a different thing: the authors assert that the conclusion follows, and they can be wrong about their own logic, so the line is firmer than any statement a point can be (1000 flips against the point cap’s 198, five times the cap) and softer than a formal step, which never bends. A reader’s what-if therefore bends adeductiveline after every premise a point holds, and aformalline never. Measured on the confusions map’s Socrates syllogism (c4: two premises at 0.99, the syllogism at 0.99, the reader’s “not mortal” at 0) through the shipped solve with the syllogism re-keyeddeductive(2026-09-08,solve_map.py --overrideand the compile’ssocrates-tiertest): at 1000 flips the syllogism gives 0.06 (0.99 to 0.93) where each premise gives 0.30 (0.99 to 0.69) and the reader’s 0 is held at 0.30; at the 200 it first had, the syllogism gave 0.24 (0.99 to 0.75), nearly as much as its premises, which is what the raise fixed. Under the map’s ownformalkey the step holds at 0.99 and each premise gives 0.32. Said plainly: on the two-premise copy thedeductivestep does bend. Its coin reads 0.93 while the premises give, which misses the condition the raise was set for (the coin above 0.95 while the premises give) by 0.02; on the single-premise copy (c2a) the same key holds at 0.986. Theformalkey holds exactly on both, 0.990. The constant stays at 1000 and the 0.02 is on Felix’s list (D161 item 3; the ladder through the compile reads 0.75 / 0.93 / 0.96 / 0.98 / 0.99 at 200 / 1000 / 2000 / 5000 / 20000 flips, pinned insocrates-tier.test.ts). - The boundary sentence.
deductiveonly where the line’s own words claim necessity, definition or elimination; a reason that would fail if the world were arranged otherwise ismechanism, however confidently the source states it. (Two blind passes over the flagship disagreed on 25 lines at exactly this seam before the sentence existed; with it they resolve by rule.) - A stated sample size raises the count and never lowers it.
# kind: empirical n=200holds a survey of two hundred at 200 flips; a line that reports one incident stays at the default 16. The floor is measured: three premise-less one-incident lines at one flip each held he-xrisk’s@containment-failsat 0.75 against the author’s own check of 0.85, and at the default within 0.013 of it (PLACEMENT_PROBE section 23 (e), 2026-09-08). - An undercut or rebuttal takes the kind of its own step, never of
the line it attacks: an undercut by mechanism is
mechanism, and “technophobia, not an inference” istestimony.
What the count changes, and what it leaves alone. The map as authored
barely moves: the flagship’s 348 keyed lines move one statement past
0.05 (@nat-abs 0.14 to 0.19) and light no new badge (2026-09-08;
under the one draw, 2026-09-24, still one, @nat-abs by 0.051, and no
badge).
What the count decides is which line gives first under a reader’s
what-if (4.6). The solve reads four counted tiers (1000, 64, 16, 4) and
the hard one, so a disagreement between empirical and testimony, or
between analogy and hope, moves no number; the word is still worth
getting right, because the reader sees it on the popup. The flagship’s 348 lines carry
their keys since 2026-09-08: 105 mechanism, 69 deductive, 65
testimony, 45 hope, 34 empirical, 30 analogy and no formal
(the book has no machine-checkable step). The map has grown to 361
lines since, every one keyed (2026-09-24: 108 mechanism, 69
deductive, 68 testimony, 48 hope, 36 empirical, 32 analogy).
Since 2026-09-26 two of them are formal: the analytic inverses
$not-feasible-not-built and $no-win-no-extinction, each true by the
meaning of its terms, at 0.99 (recount that day: 106 mechanism, 69
testimony, 66 deductive, 48 hope, 35 empirical, 31 analogy, 2
formal). Since 2026-09-27 a third line is formal,
$harmless-no-extinction, one of the two inverses on the title claim
(7.7, D168).
The rubric’s record, with the
two-pass agreement (kappa 0.65) and the third reader on the ties, is
ideas/plans/semantics-antecedent-pull-2026-09-06.md section 17.
4.2 The unpriced ground, and when to price it
A line says nothing about its own premise. If the only line in a map is
$imp 0.8 @c | @a, the solved P(@a) is 0.500 and P(@c) is 0.700: the
network gives @a even odds, the value it gives any claim nothing
speaks to, and in the half of the worlds where @a holds the line
brings @c to 0.9 (the floor of 4.1). A premise on the negated side,
$e 0.9 @c | ~@a, stays at 0.500 as well (both measured 2026-09-24
with solve_map.py at its default; the first is
examples/toys/a10-t2.argmap, group ::free, in the app the Semantics
audit map audit-reference, group R1).
That 0.500 is a placeholder. Nothing in the map says what the premise is worth, so the solve leaves it at the network’s fill and lets the lines that use it settle it, which is rarely what an author means by a premise. The checkers name every such statement with the info note I7, an unpriced ground (10.1), and the fixes are ordinary authoring:
- Author a value on the premise. On a frontier root this is the whole fix: a point or a pair at the source’s register (4.7).
- If the source asserts the premise as a claim of its own and you want it discussable as one, give it an attributed premise-less line at that register (the direct-assertion pattern, 7.13).
- If the premise was only ever part of the step, fold it into the line that uses it and re-elicit that line’s strength.
A converse is a different thing, and still worth writing. If the source
also asserts “and otherwise not” (“nobody builds it, it kills nobody”), write it as a
second line on the other side of the premise, ~@c | ~@a. It speaks to
the worlds the first line leaves open, the ones where the premise
fails, so it moves the conclusion and leaves the premise alone: with a
0.8 converse beside the 0.8 line, @c reads 0.500 and @a stays at
0.500 (the Concepts map c1-drift-tax, block 3). It also keeps the
claim on the map, where a reader can argue with it. The flagship map’s
$no-doom-otherwise is the worked example (7.7).
What does move a premise is information about what follows from it:
- A confirmed consequence raises it: pin
@cat 0.9 beside the lone line and@areads 0.593, and each further confirmed consequence raises it again (the fork gallery’sf-consilience, MATH §5). - A refuted consequence lowers it (modus tollens;
f-refutation). - A support and an objection on the same premise whose strengths sum
past one make their shared case rarer: 0.8 against 0.3 reads
@aat 0.466, where 0.8 against 0.15, which fit, leave it at 0.500 (4.4).
Each is the map’s own inference about the premise, and the residual rule (4.3) says to leave it standing.
Until 2026-09-08 this section taught the drift tax: under the
retired uniform reference a lone 0.8 line dragged its free antecedent
to 1/(1+2^0.8) = 0.365, a line on the negated premise pushed it up to
0.651, a confirmed consequence lowered it (0.464 where the shipped
solve reads 0.593), and the remedies were ways to cancel that pull. The
network reference D161 ships has no such pull; MATH §4.4 keeps the
record, and solve_map.py --reference d36 reproduces every number of
it (the same scratch toys, 2026-09-24).
4.3 Statement values and the residual authoring rule
A statement’s authored value is a floor-style constraint, and the single most important discipline in the whole system applies to it:
Author only the evidence for or against a statement that is not already contained in the rest of the map.
- Frontier roots (statements with no incoming evidence in the map) keep their authored values. Their number is the map’s interface to everything unmapped; that is what roots are for.
- Derived statements (concluded into by mapped evidence) should normally carry no authored value. Their probability is the output of the solve. If you author one anyway, you are counting the mapped support twice.
- If you disagree with what the solve delivers for a derived statement,
you have three honest moves, in order: fix the argument (structure or
strengths); add the missing evidence as a new, named line (a
premise-less evidence is fine, but it must say what the evidence is);
or record your number as a check credence,
# check: p, and let the displayed badge show the disagreement.
The check credence is the designated home for “all things considered I believe 0.85 even though the mapped argument delivers 0.48”. It is displayed, compared, and never constrains the solve. Wanting to force a derived statement to a number is precisely the situation the rule exists to catch.
One explicit anti-pattern: a statement line carrying both an authored
value and a # check: comment. On a concluded-into statement the pin
double-counts the mapped support and, worse, fights the very evidence
you authored against it (the pin holds the solved value where the
counter-evidence should have moved it), while the check silently
disagrees with the pin. Concluded-into statements take a check or
nothing; only frontier roots take pins. The lint and the editor both
flag a bare point beside a check as W23 (since 2026-08-08, narrowed
2026-08-22 to the bare point), on any statement line, because it is
wrong wherever it appears: one extraction wrote both on 28 nodes, and
the only visible symptom was that the gaps looked suspiciously small. A
bare point beside a strengthed line that concludes into the statement
is W25 (2026-08-22), whether or not a check sits beside it. An
explicit residual pair beside a check (0.6?/0? with # check: 0.9)
is silent: that is the intended two-slot shape, the pair being the
unargued remainder and the check the total (4.7).
The rule is topological and does not change inside refinement boxes. A hinge statement that sibling lines inside a refinement conclude into is a concluded-into statement like any other: check, not pin, even when the source asserts it at a clear register and the mapped internal support delivers less. That under-delivery is an audit finding about the source, not a display problem to pin away. When the source asserts the hinge directly, over and above the arguments it gives for it, that assertion is itself evidence and has a named home: a premise-less, attributed evidence line at the source’s register (the direct-assertion pattern, 7.13). It accumulates with the argued routes instead of clamping over them, and it is visible and criticizable in a way a pin never is. Under v0.3, a floor pair is the interval-shaped variant.
The distinction that keeps this straight: a statement’s own indented block explicates its number (the block refines the implicit factor asserting the marginal, and the statement keeps it, 3.4); sibling lines concluding into the statement replace it (the value becomes the solve’s output, and the author’s number moves to a check).
Also: inference through the map is not evidence you authored. If an
evidence about @a -> @c moves the solved P(@a) (a confirmed or a
refuted consequence, 4.2), do not “correct” @a’s authored value for it; that effect is already
contained in the map.
A root’s number has a second axis too (D161, 2026-09-08): how firmly
the solve holds it under a reader’s what-if comes from the width the
author left open, an interval 0.7/0.1 about 8 flips and a point about
200, the cap. Nothing is elicited for it; the width you write for the
register (4.7.4) is the firmness. Section 4.7.2 gives the rule.
4.4 Independence, and what to do when it fails
Separate evidence lines are treated as independent mechanisms; their premise-less masses accumulate like independent reasons (noisy-OR). That is what makes two convergent lines mean “two independent reasons”. When the grounds actually overlap, independence double-counts. Three repairs, in increasing order of structure, follow the next paragraph.
Independence holds among the lines on one side. A line for a statement and a line against it that apply to the same case are read together (since 2026-09-23, MATH §3.8): they are shares of one population of cases, so they never fire in the same world, each keeps the share its author gave it, and independence is what gives where the numbers do not fit. A granted objection therefore caps what the supports can claim in the cases where it applies, however many supports there are: five 0.9 supports beside a 0.15 objection read 0.90 in the case where all their grounds hold, where weighing them as independent evidence read 1.00 (measured 2026-09-23 on a seven-statement toy with every ground free). To move an objection, answer it with an undercut (4.5), lower it, or doubt its grounds. Where the strongest support and the strongest objection sum past one, the overlap is a contradiction that makes the case rarer, and the pressure lands on the grounds that produce it. If the two lines really describe different cases, name the statement that separates them, and they stop meeting.
The three repairs for overlapping grounds:
- Merge the lines into one evidence if they are really one argument.
- Name the shared source as a statement and condition both lines on it; the dependence is then authored in the world layer where it belongs.
- Complementary partition: make an “even if” explicit by conjoining
the negation of the other route, as in
$mwb-time ... | @wont-solve-in-time AND ~@align-hard(the two routes then partition the worlds instead of overlapping).
One consequence of independence is worth meeting before it surprises
you. Two independent reasons for one conclusion stop being independent
once the conclusion is known: with the effect settled and one cause
observed, the other cause is less likely, because the effect is already
accounted for. That is explaining away, and the solve shows it, as
any probabilistic network does. D161 declares it as a divergence from
the conditional-logic literature’s syntax-splitting postulate, which
would hold the other cause where it was; the map takes the Bayesian
side. The exhibit is examples/toys/f-explaining-away.argmap (in the app,
the Semantics audit map audit-forks, group F6): two
causes each bring the effect at 0.9; with only the effect observed both
causes read 0.573, and observing one of them as well drops the other
to 0.515 (measured 2026-09-08 with solve_map.py at its default on the
toy, the same numbers the kernel is pinned to in the solver-compile toys
test; the record’s hard-mode figures are 0.573 and 0.512). Nothing
needs authoring around it. A reader who pins one
cause of a known effect in what-if mode and watches the other cause
fall is seeing this.
Two boundary clarifications. First, the discipline applies to lines converging on the same conclusion; one statement legitimately feeds premises of several different conclusions, and that needs no declaration. Second, genuinely independent routes stacking a hub high (noisy-OR takes four 0.85 routes past 0.99) is not by itself an error: if the source really asserts four independent sufficient reasons, the high number is what the source’s own logic delivers, and a lower check credence on the hub turns the difference into a visible audit finding (the source claims less than its own arguments compound to). Before accepting that reading, check whether the routes share an unnamed latent (repair 2); several “distinct” failure modes of one mechanism usually do.
4.5 Undercut strength
An undercut (3.6) carries the defeater’s operative rate: granted the grounds, how often does the targeted inference actually fail? The grounds’ own plausibility is carried by the guard statements, so do not pre-discount the undercut for it. Likewise, do not pre-discount a defeater because a response to it exists; author the response as its own undercut of the undercut and let the graph do the discounting. q’ = 0 is inert, q’ = 1 eliminates the target in context, values between interpolate.
An undercut does not, however, push its own conclusion. This
paragraph said the opposite until 2026-07-27, when the skill-only
sufficiency eval authored a cluster on the strength of it and produced a
0.61 statement gap; the claim is wrong and the correction matters for
authoring. An undercut-shaped line compiles as a pure inhibitor of
its target: per the factored-A compile (SOLVER_SEMANTICS §1.2),
“inhibitors carry no zero-set of their own”, so the negated conclusion
the line names receives no independent floor from it. Measured under
the shipped solve (2026-09-24, solve_map.py at its default): an
undercut whose target is unstrengthed leaves its conclusion at 0.500,
exactly as if the line were absent. And an answer to an objection, an
undercut of a rebuttal, recovers the claim toward the value it would
hold with the objection absent and never past it. The flagship’s shape,
in four statements:
---
title: an objection and its answer
argmap-format-version: 0.3
---
@s [S] 0.8: the support's ground
@w [W]: the objection's ground, left open
@g [G] 0.85: the answer's ground
@c [C]: the claim
$sup [S supports C] 0.9 @c | @s:
$obj [W argues against C] 0.2 ~@c | @w:
$ans [G answers the objection] 0.9 @c | @g AND $obj:
C reads 0.859 with the support alone, 0.823 once the objection stands
unanswered, and 0.840, 0.852 and 0.855 with the answer at 0.5, 0.9 and
1; even at 1 the answer leaves the objection standing in the 15 percent
of cases where its own ground fails. Take the support away and the
objection with its answer reads 0.488: above the objection alone
(0.450), below the 0.500 of a claim nothing speaks to. The answer
revives the worlds the objection killed and delivers nothing of its
own. The file as shown is the toy examples/toys/u-grounds.argmap,
group ::guard-sup with its ids unsuffixed, and the supportless one its
group ::guard (in the app, the Semantics audit map audit-undercuts,
group U3); the other readings are the same lines with one changed or
removed. T13’s reinstatement sweep (u-t13,
audit-undercuts U2) reads the same way on the inference itself. The
figures this paragraph quoted until 2026-09-24 (0.500 to 0.866 against
a rebuttal-free 0.898) were measured in July under the retired uniform
reference.
The authoring consequence: when the source both raises an objection to
an inference and asserts the fact that objection rests on, the
undercut carries only the first. If you want the fact to bear on the
claim as well, give it its own ordinary evidence line beside the
undercut. That is not double-counting (the inhibitor acts on the
inference, the plain line acts on the claim), and without it the fact
the source actually reports is silently absent from the solve. The
same holds for an answer’s ground: wired as its own 0.9 line into C
beside a guardless answer, the ground brings C to 0.879 where the same
ground inside the answer’s guard gave 0.488 (u-grounds, ::split
against ::guard, measured the same day).
Which objections are undercuts at all (D170, 2026-09-28). An
undercut says a step is unreliable; where it holds, it says nothing about
which way the claim goes, so a reader who wins an undercut outright is
left with cases no link speaks to, and those count as even odds. Before
wiring an objection, ask what is true instead if it is right. If the
answer is that the step cannot be trusted, it is an undercut. If the
answer names a claim that is false, it is an ordinary line into that
claim’s negation, and the answers that say its inference fails stay
undercuts of it. On the flagship, “experts disagree” stays an undercut of
$link-fine, while “if building killed the builders, they would stop”
and “nothing so far has ended us” are lines against @mis-ext. Won
outright, each now takes the title claim to 0.10 through
$harmless-no-extinction; as undercuts, the first left it at 0.37
(measured 2026-09-28 with solve_map.py at its default).
Readers arriving from probabilistic conditional logic: the translation rules under 3.6 say when a low conditional is an opposed line and when a subclass exception needs this undercut plus a rebuttal.
4.6 Solved values, tension, and the what-if
The display is computed-first: the solved value is the primary number, the authored values are the claims it was asked to honour, and tension is the distance from the solved value to what the author wrote. Since D161 that distance is measured against the whole of what was written: a point, an interval’s two bounds, a step’s strength as a lower bound. A tint means “the map could not honour this number, by this much”; the solve settling somewhere inside an interval the author left open is a readout without colour (2.2 has the rule, the 0.01 tolerance and the flagship’s counts). A tension badge is information about the argument, and it stays: the flagship’s widest, the headline 0.084 short of its check, is an audit finding about the source (2.2).
Authored 0 and 1. Since D161 a point is held at the point cap, about 200 flips, so an authored 0 or 1 no longer deletes possible worlds (it did under D36: “world-killers”, SOLVER_SEMANTICS P3). It still claims a certainty the source rarely states, and it is held no more firmly than 0.97 is. If you mean “very confident”, write 0.97, or give the interval as a v0.3 pair. The honest wide statement is cheap; the false point is not.
The what-if is revision. In what-if mode a reader moves a claim’s
number. What the map does with it was ruled on 2026-09-08 (D161, after
the two-verb discussion in
ideas/plans/semantics-antecedent-pull-2026-09-06.md section 15), and
it is one rule with two parts:
- The reader’s number replaces the author’s on that claim (the author’s point or interval there is set aside for the duration) and enters as a point at the cap: your number is a claim like the author’s, as firm as a point (198 flips, the cap). A reader’s 0 or 1 is therefore a very firm claim, never a fact.
- Everything else the author wrote stays in force, each number at its own firmness, and the map re-solves. The question answered is: if this number were as the reader says, which of the author’s other numbers gives, and by how much.
Two cases follow, and both are worth knowing. The numbers are the
flagship’s under the ruled counts (width counts on its intervals, kind
counts on its lines with deductive at 1000, the reader’s row at the
cap; re-measured 2026-09-23 on the map as it now stands with
solve_map.py under the shipped solve, the one draw of MATH §3.8
included; the 2026-09-08 figures this section quoted before came from an
earlier state of the map and the rule before the one draw):
On a root, the map follows you forward. @llm-nice (the authors’
0.85/0.05 that current systems seem nice) moved to 0.1: the map’s
answer is @nat-abs 0.20 to 0.11 and nothing else past 0.01; no line
bends further and nothing new colours. The reader changed a premise the
author argued nothing for, so there is nothing to give; the map
propagates. A better-connected root shows the same at scale:
@instrumental-convergence to 0.2 moves eleven statements past 0.01
(@incorrigible 0.83 to 0.57, @not-preserved 0.89 to 0.77), lights one
of the book’s own check badges downstream (@not-preserved at -0.18: a
conclusion the reader’s premise no longer delivers) and bends three
hope lines, the hopes that argue against that conclusion
($c5-selfish-obj 0.70 to 0.63, $c5-digital-obj to 0.67,
$c5-leave-obj to 0.68). Every tint sits downstream of the edit.
On a conclusion, the author’s case retreats where it is softest.
@everyone-dies (the headline; no authored number, check 0.88..0.98
when these figures were measured, re-read to 0.75..0.95 on 2026-09-25)
moved to 0: the map meets the reader (0.0005), bends no line past 0.05,
and gives on the free premises upstream, @if-built 0.88 to 0.53 (its
check badge lights at -0.32), @asi-soon 0.89 to 0.70, @mis-ext 0.97
to 0.88. Read it as “then it was not built, or not soon”: the book’s
case is softest at its premises, because they are derived claims nobody
pinned. Pin those too (the record’s P6 set: built, misaligned when
built, the misalignment extreme, the link fine, and still nobody dies)
and the retreat has to land on the lines and on the authors’ flat
assertions: the analogy response $resp-exp 0.90 to 0.63, the
testimony undercut $uc-experts 0.30 to 0.54, the mechanism
objection $c12dr-obj 0.15 to 0.21, a second analogy
$uc-precedented 0.30 to 0.35, and two 18-flip assertions leave their
intervals (@precedented 0.11 to 0.25 against 0.05..0.15,
@immature-field 0.89 to 0.83 against 0.85..0.95); no deductive
line moves, and no hope line does either, because none sits on the
path: the count orders what gives among the numbers the reader’s claim
reaches, which is why a 64-flip mechanism bends here while the 4-flip
hopes stand. That list is what “the book’s case holds least firmly”
means, and the Most moved panel shows it with each line’s kind. Since
2026-09-28 the two objections that deny @mis-ext argue against it
instead of doubting the step (D170, section 4.5), so the step keeps one
doubt to bend, and the same set’s retreat lands on $uc-experts, 0.30
to 0.98, and $resp-exp, 0.90 to 0.27, alone.
The refused residual. The reader’s number is met to within 0.002 in
every case above (through the shipped compile and kernel the P6 set’s
points read 0.0020, 0.9981 and 0.9982 for @everyone-dies, @if-built
and @mis-ext, D161 item 4; @llm-nice at 0.1 reads 0.101, both
re-measured 2026-09-08; since 2026-09-28 the P6 points read 0.0044,
0.9956 and 0.9956, one undercut being left to bend), and that is the usual outcome, because 198
flips outrank everything on the flagship except its 69 deductive
lines.
Where the map cannot meet it, the adjustment row says by how much. On
the confusions map’s Socrates syllogism (c4: two premises at 0.99 and
a formal step at 0.99) a reader’s “not mortal” at 0 is held at 0.32,
the step holding at 0.99 and each premise giving 0.32; re-keyed
deductive, the step gives 0.06 and the reader’s 0 is held at 0.30
(4.1, rule 1; both re-measured 2026-09-08 through the shipped solve).
The residual is the honest answer, and it is tinted only past the 0.01
tolerance.
The other reading of a what-if, “suppose it turned out that way”, is
the conditioning verb: the reader’s number as a fact about the world,
the author’s numbers updated by Bayes. It stays a bench readout
(solve_map.py --condition, 2.3). Under conditioning the author’s pins
would yield; under revision they hold at their firmness and the reader
sees which inference they thereby reject. The reason for the choice
(recorded in semantics-antecedent-pull-2026-09-06.md section 15 item
12): showing a reader where an argument’s inconsistencies lie is the
point of the tool, and conditioning hides that by moving to the
unlikely world without showing how surprising it was. One more effect to expect in
what-if mode: pinning one cause of a known effect lowers the other
causes (explaining away, 4.4).
4.7 The epistemic delicacies: residual, point, pair, check, derive
The rules above keep a map lint-clean. The rules in this section keep a
lint-clean map from counting one consideration twice, which is the error
class no lint can see in full and the one an extractor has to get right
in a single pass. They were settled in August 2026 (DECISIONS.md D152,
the register to pair table in AUTHORING_NOTES 2026-08-23, the check
intervals of the same day) and each of them has a five-line measurement
behind it. Those measurements are the toy battery in examples/toys/; a
reader-facing selection of them ships in the picker’s Semantics audit group
(examples/audit/, generated from the toys with the measured answers in the
group heads). Read the toy, then the rule; the rule is what the number says.
Two measurement dates sit in this section, and they are marked. The
T1 to T4 illustration in 4.7.2 was re-measured on 2026-09-08 under the
network reference (D161, the shipped default since that day) with
solve_map.py at its defaults, as authored and under a reader’s
what-if; the toys’ README carries the same numbers in a dated block.
The pictures from 4.7.5 on were re-measured on 2026-09-08 the same
way (solve_map.py at its defaults; the README’s second table), and
each names the retired D36 figure it replaced beside the network
reference’s. The rules stand under both references, because they are
about what a number is evidence for, which no reference changes; what
the reference changes is where the double count becomes visible, and
4.7.2 says where. One picture did not survive the change and says so:
the ridge price of a p = 1 terminology statement (4.7.6, T7) was the
uniform reference’s, and under the network reference the statement
holds at 1.000 and the junction pays nothing.
4.7.1 The residual rule, stated once more
A statement’s own number is evidence not already in the map. Two cases follow:
- A frontier root (a statement no strengthed line concludes into) keeps its number whole. Nothing in the map argues for it, so its number is the entire interface to the unmapped world, and the all-things-considered reading is the right one.
- An interior statement (at least one strengthed line concludes into it) may carry a number only for the part of the author’s credence that the mapped lines do not trace: the unargued remainder, which the reader-facing vocabulary calls the author’s tacit grounds. The author’s total belongs in the check (4.7.2).
The asymmetry is the whole rule. A root’s number and an interior statement’s number are different kinds of object, and the mistake that produced the 2026-08-03 incident (AUTHORING_NOTES) was writing the same all-things-considered figure into both slots.
4.7.2 Point, pair, check: three slots on one statement
A statement line has three places a number can sit, and they mean three different things.
- The point,
@s 0.9?. Since D36 a bare point is read as the degenerate pair0.9?/0.1?, a zero-width interval: the members sum to 1 and P(s) is held at 0.9. Since D161 it is held at the point cap: a point counts as an interval with silent mass 0.01, 198 flips (about 200), so no two points can make the solve infeasible, and a point may solve up to about 0.005 off its value, inside the tint’s tolerance. The cap sets the firmness only; the target stays at 0.9 and no interval is written for it. (Under D36 the point was a spring at weight 400 that leaked about 1/400.) - The pair,
@s 0.6?/0?(v0.3, D53). The author’s direct evidence about the statement: a floor of 0.6 for it and nothing against it, so P(s) is bounded to [0.6, 1]. The map’s own inference then selects within the interval: every line the statement feeds or is fed by pushes the solved value around inside it, and the reference picks the point when nothing pushes (under D36 the balancing prior’s maximum-entropy point; under D161 the map’s own network, where the lines put it, and on the flagship every one of the 224 intervals is met inside its bounds, 111 of them more than 0.05 above the floor, 2.2). Inference through the map never counts against the pair; the pair is read as direct evidence and the width is the author’s honest spread. Since D161 the width is also the firmness: by Walley’s imprecise Dirichlet model with prior strength 2, an interval with silent mass m is held at 2 (1 - m) / m flips, so0.6?/0?(silent 0.4) is 3 flips,0.7/0.1is 8,0.85?/0.05?is 18 and the cap’s 0.01 is 198. A wide interval is an honest spread and an easy give under a reader’s what-if; that is one fact, said twice. - The check,
# check: 0.85..0.95. The author’s total, all things considered, as an interval at the register’s width (4.7.4), or a bare point as the zero-width case. It never constrains the solve, so it has no firmness either: under a reader’s what-if a check-only statement is the first thing to move, and its badge says by how much (4.6’s@if-built). The badge a reader sees is the signed distance from the solved value to the interval, zero inside it. A badge that stays is a finding about the argument; closing one by moving a number is the move the whole discipline forbids.
The three toys T1, T2 and T3 put the three slots on the same five-line
map. @want is a frontier root, @subvert is the statement whose
authored 0.9 silently contains “because it wants things”, and
@resists reads @subvert.
---
title: T1 pin kept beside a mapped argument
argmap-format-version: 0.3
---
@want [The system wants things] 0.9?: a frontier root
@subvert [It routes around oversight] 0.9?: the author's all-things-considered 0.9, kept as a PIN
$sub-ev [wanting means routing around obstacles] 0.9? @subvert | @want:
@resists [It resists being corrected] : # check: 0.7
$res-ev [subversion implies resisting correction] 0.85? @resists | @subvert:
The numbers below are the network reference’s, measured 2026-09-08
with solve_map.py at the shipped defaults, first as authored and then
under one reader’s what-if (--override want=0.1: the reader’s number
replacing the author’s on the root, 4.6).
T1 is the shape to avoid, and the lint says so (W25 on @subvert): the
point sits beside a strengthed line that argues for the same statement.
As authored it solves @want 0.899, @subvert 0.900, @resists
0.883. The wanting consideration is counted twice here, once inside
the 0.9 and once through $sub-ev, and the as-authored solve shows no
sign of it, because a pin mostly restates itself.
T2 derives the statement: the same file with @subvert’s head line
replaced by @subvert [It routes around oversight] : # check: 0.9. As
authored it solves @want 0.899, @subvert 0.905, @resists 0.884,
the same picture as T1 to within 0.005. Under the network reference
the one line the map wires delivers the author’s total by itself: a
0.9 line under a 0.9 premise fills its conclusion at 0.9 × 0.95 + 0.1 ×
0.5 = 0.905 (the F1 fill), so the check is met and no badge lights.
That is the honest state too. The badge is the distance between what
the mapped lines deliver and what the author holds, and here they
agree; it lights the moment they stop agreeing, as the what-if shows
next. (Under the retired counting reference T2 read @subvert 0.868
against the check and @resists 0.866 against T1’s 0.879, a badge and
a drop that earlier versions of this section presented as the double
count made visible; both were that reference’s antecedent pull and
went with it, D161 item 1.)
Where the double count shows under the network reference is the
what-if. With the root @want set to 0.1, T2 follows its premise down:
@subvert 0.545 against its check of 0.9, the badge lit at -0.36, and
@resists 0.732. T1 does not move: @subvert 0.899 and @resists
0.882, the pin holding at its 198 flips against a 16-flip line, exactly
as it held against the map. A pinned interior statement is deaf to its
own premise, and that is what counting a consideration twice means: a
reader who doubts the wanting is told it makes no difference to the
routing the author derived from the wanting.
T3 adds a residual floor pair beside the check: @subvert [It routes
around oversight] 0.6?/0?: # check: 0.9. As authored it solves
@subvert 0.896, @resists 0.881, and the lint is silent, because a
non-coinciding pair beside a check is exactly the two-slot shape D36
item 3 describes. Under the same what-if it reads @subvert 0.840 and
@resists 0.857: the floor absorbs most of the reader’s move (the map
now explains @subvert through the floor instead of through @want),
which is fine when the 0.6 is the unargued remainder the text licenses
and is the pin again with extra steps when it was read off the total.
T3’s hazard is elicitation; the machinery is sound.
T4 is the legitimate interim state: the pin kept, the relation drawn,
the strength left off ($sub-ev [..] @subvert | @want:). An
unstrengthed line compiles inert, so T4 solves as the pin alone (@want
0.899, @subvert 0.899, @resists 0.882), the lint lists the line
under I3, and nothing is asserted twice. Against T1 the strengthed edge
beside the pin is worth 0.001 as authored and nothing under the what-if
(T1 and T4 both answer @subvert 0.899, @resists 0.882 at @want
0.1): next to a point at the cap, a 16-flip line adds nothing the pin
had not already asserted. (Under the counting reference the lift read
+0.006 and +0.002 and stood here as the double count visible even in a
toy; under the network reference the count makes the pin outrank the
line outright, and the visible sign is the refused what-if above.)
4.7.3 Deriving a root (D152)
To derive a root is to give a pinned root its first incoming strengthed line. One test decides whether the line may be written, and one obligation follows once it is.
The inference test. The line is written when it carries an
inference (premise to conclusion) and withheld when it would only
restate, or co-refer to, the same episode or proposition at a second
dock. A restatement wired as inference counts one consideration twice.
The flagship’s precedent is @psychosis and @c13ws-retrain: one
episode family at two docks, cross-referenced in comments and never
joined by a line. Source faithfulness is no part of the test: a real
inference the source omits is a comprehensiveness defect of the source,
and it is mapped, with a comment marking it as the mapmaker’s.
What does not count against a derivation. “A heavily consumed root,
once derived, moves everything downstream” is true and is no reason to
withhold the line: the movement is the coherence audit working.
Completion is always licensed; if a correct completion swings the
conclusions wildly, that refutes the semantics or the elicitation and
never the completion. Measured on the case that settled it: deriving
@steering-finds-subversion from the two grounds the book itself gives
(| @goals AND @deep-gears) moved that statement 0.893 to 0.628,
@goals 0.842 to 0.712 and @incorrigible 0.821 to 0.695 (the two-floor
pass, AUTHORING_NOTES 2026-08-23, under the retired uniform reference,
whose antecedent pull amplified every such move). Under the shipped
solve the same derivation moves it 0.899 to 0.843 and @incorrigible
0.856 to 0.834, and leaves @goals at 0.851 (the map as it stands
against a copy with the pin restored and the line removed,
solve_map.py at its default, 2026-09-24). Either way it is a finding
about the book (its flat 0.9 outruns its own two grounds), and the map
now shows it.
The elicitation obligation. Once a line sits under the root, its authored point is no longer a residual. The point moves to the check, at its register’s interval, and the residual slot is empty by default for every register: a floor read off the register is read off the total, because the register describes the all-things-considered credence and the text never gives the split. Two text-licensed exceptions, applied row by row:
- The passage names a second ground the map does not yet wire: a floor at that ground’s own register, with a gloss naming the passage, as the interim encoding until the line is written.
- The passage says the listed grounds are a subset (“to name a
few”): a remainder at the fallback row
0.2?/0?, with the quote in the gloss. Small enough never to close a badge; nonzero so the acknowledged mass shows.
In practice the derivation is one head-line edit plus one new line:
# fragment - not standalone (flagship, anchor "$subversion-ev"; before and after the 2026-08-23 derivation)
@steering-finds-subversion [A goal-directed system seeks ways to subvert whatever limits it] 0.9?: shared ground (reused by Ch5, Ch11) # C7 flat: the state before
@steering-finds-subversion [A goal-directed system seeks ways to subvert whatever limits it]: shared ground (reused by Ch5, Ch11) # check: 0.85..0.95 (C7 flat; derived via $subversion-ev)
$subversion-ev [wanting routes around obstacles the gears can find] 0.9? @steering-finds-subversion | @goals AND @deep-gears: the two grounds the authors give
(The fragment shows both head lines for the comparison; a real file carries one of them.) Before writing the edge, grep the map for the root’s id: the reason it was left unwired is usually written down in a comment somewhere, and a derived root’s comment has to be rewritten along with its number.
4.7.4 The register to pair table, in brief
The source states neither points nor intervals; it states registers (flat repetition, “probably”, “our best guess”, “nobody knows”). The D39 rubric mapped registers to points; the 2026-08-23 table adds the pair column, pre-registered before any root was touched. Its rules:
- Centring. A pair’s midpoint (1 + p⁺ − p⁻)/2 equals the rubric point, and the register fixes the width: p⁺ = point − w/2, p⁻ = 1 − point − w/2. Widths are rubric constants: 0.05 for the strongest categorical class, 0.10 for flat assertion, 0.20 for “by default” and “best guess”, 0.30 for “we expect” and “could well”, 0.40 for the weakest hedge, 0.80 for a refusal. Authored mass is linear probability; log-odds widening was only ever a sweep’s parametrisation.
-
The rows, point to pair, for a root that stays a root:
register (flagship class) point pair repeated categorical, anti-hyperbole (C1) 0.97 0.95?/0?categorical deontic, “insane gamble” (N1, N3, RS1) 0.93 0.9?/0?title register, C2 0.93 0.88?/0.02?flat unhedged assertion (C7, P-FACT) 0.90 0.85?/0.05?“by default”, “our best guess” (C3) 0.85 0.75?/0.05?RS3 0.80 0.7?/0.1?“we expect”, a granted point (C4, P-GRANT) 0.75 0.6?/0.1?“could well” (C5) 0.70 0.55?/0.15?weakest hedge (RS6) 0.60 0.4?/0.2?role defaults 0.8 / 0.7 / 0.6 0.65?/0.05?,0.55?/0.15?,0.45?/0.25?explicit refusal, “nobody knows” (COIN) 0.50 0.1?/0.1?plus a gloss sentencethe objection the source denies at class X (P-CONTEST) 1 − X the mirror of X’s pair - One-sided categoricals. Where the source leaves no room against
(
0.9?/0?,0.95?/0?), the midpoint sits above the point, and that is the reference’s fill (the balancing prior’s under D36, the network’s under D161), never a number the table asserts. - The P-CONTEST mirror. An objection the source denies at class X
takes X’s pair reflected (
0.05?/0.85?for a flat denial,0.1?/0.6?for a hedged one), never0/p: omitting the floor would leave nothing for the objection, and “we don’t think so” is weaker than “impossible”. - Refusals are wide pairs with a gloss. “Nobody knows” where the
statement is the unknown becomes
0.1?/0.1?and a gloss sentence naming the passage, in one shared shape: what the passage does (“the page calls the question open”), then “the wide pair carries that, and its midpoint is nobody’s belief”. The interval is the message. - Per-node asymmetric pairs are written where the passage states both directions (an inner flat claim under an outer hedge, a concession, “routes against their own”): from the text, with the rationale in a trailing comment, and the midpoint may differ from the class point. The class row is the default where the text gives one direction.
- Riding rules.
?on every member, zero members included; a gloss sentence whenever the total width exceeds 0.30; the class token in the trailing comment of every pin; nothing on an evidence line moves. - The width is the count (D161, 2026-09-08). The same widths fix
how firmly the solve holds the pair under a reader’s what-if, by
Walley’s rule n = 2 (1 - m) / m on the silent mass m: width 0.05 is
38 flips, 0.10 is 18, 0.20 is 8, 0.30 is 4.7, 0.40 is 3, the
refusal’s 0.80 is 0.5, and a point’s capped 0.01 is 198. On the
flagship the median pin is 18 flips and the firmest 38, below a
mechanismline’s 64: under revision the authors’ assertions give before their mechanisms, which is the completion principle’s order and the story 4.6 tells.
One proviso, stated so it stops being rediscovered: a root’s register-read confidence is a posterior that may already discount what the author believes downstream, and reading it as direct evidence lets the solve apply modus tollens a second time. There is no operational alternative (the text never gives the un-discounted figure), the magnitude at these widths is small and measured (at most 0.043 per statement in the widening sweep, mean 0.006 on the centred table), and the compensation is to show the interval and the band rather than present a midpoint as the author’s number.
4.7.5 Shared considerations
Separate lines are combined as if independent (4.4). The only way the map can say two lines co-vary is a shared statement they both cite, so whenever two lines rest on one consideration, name it as a statement and condition both on it. The sign of the effect tells you whether you had it right. T5 and T5b put one consideration into an AND junction twice, first through a pin and then through the shared root:
---
title: T5b the same junction with both routes derived
argmap-format-version: 0.3
---
@want [The system wants things] 0.9?: a frontier root
@subvert [It routes around oversight] : derived from the same root the other route uses # check: 0.9
$sub-ev [wanting means routing around obstacles] 0.9? @subvert | @want:
@grabs [It grabs resources] : # check: 0.9
$grab-ev [wanting implies grabbing] 0.9? @grabs | @want:
@danger [It is dangerous] : # check: 0.9
$danger-ev [subversion and grabbing together] 0.9? @danger | @subvert AND @grabs: the wanting consideration enters twice here
With @subvert pinned at 0.9 instead (T5) the conjunction solves
@danger 0.866; derived as above it solves 0.876 (network reference,
2026-09-08; 0.843 and 0.850 under D36). The conjunction
rises once its two premises are correlated through the shared root,
because independence understates a conjunction of positively correlated
premises, and the pin had asserted independence.
T8 and T8b are the same fingerprint with both consumers present: two
study findings that share one methodology statement @method (T8) or
rest on separate @method-a / @method-b roots (T8b), each read by an
AND consumer and an OR consumer.
---
title: T8 two lines sharing a premise co-vary through it
argmap-format-version: 0.3
---
@method [The shared methodology is sound] 0.7?: the common cause, named as a statement
@a [Study A's finding holds] : # check: 0.65
$a-ev [A's report, given the method] 0.9? @a | @method:
@b [Study B's finding holds] : # check: 0.65
$b-ev [B's report, given the method] 0.9? @b | @method:
@both [The effect is real, both needed] : # check: 0.6
$both-ev [two findings together] 0.9? @both | @a AND @b:
@either [The effect is real, either suffices] : # check: 0.8
$either-ev [either finding suffices] 0.9? @either | @a OR @b:
The two findings do not move between the maps (@a and @b 0.815
in both; 0.750 against 0.751 under D36). The consumers move in opposite
directions: the AND rises (@both 0.799 separate to 0.818 shared;
0.751 to 0.782 under D36) and the OR falls (@either 0.935 to
0.915; 0.918 to 0.886 under D36), because findings that share a
methodology stand or fall together, which helps a conjunction and costs
a disjunction. argmap-query shared-cause lists every such overlap in a
map, one row per shared statement; each row is either the deliberate
world-layer device or an overlap to factor out, and the remedy is the
same either way: name it.
One shape the static audit cannot see: a statement conjoined with its
own derivative. The flagship had $wst-race conclude from
@race-dynamics AND @one-cavalier-suffices while $ocs-ev derived the
second premise from the first, which counts the race consideration twice
through one junction (the covariance probe’s single hit above its floor,
+0.029 of manufactured co-movement). The fix is to drop the duplicate
premise, or to re-elicit the line conditionally on the derivative alone.
Trace each premise’s ancestry before writing an AND.
4.7.6 Definitional versus substantive bundling
A node that bundles a definitional part with a substantive part
is two variables in one slot. “A goal-directed system seeks ways to
subvert whatever limits it” contains “routing around obstacles is what
wanting means” (a definition) and “oversight is one of the obstacles”
(the claim). Two remedies: split the node into one proposition per
statement, or derive it from both halves, one line per ground (the
$subversion-ev shape above). No lint sees this; the authoring rule is
the tool.
A definition that does inferential work encodes as p = 1 evidence lines, and a biconditional as two of them, written as the converse pair: the forward line and its contrapositive, both concluding into the defined term (T6, respelled 2026-08-23):
---
title: T6 a definition encoded as a converse pair
argmap-format-version: 0.3
---
@steers [It steers toward outcomes] 0.9?:
@routes [It routes around obstacles] 0.85?:
@want [It wants things] : wanting, by definition, is steering plus routing # check: 0.8
$def-fwd [definition, forward] 1 @want | @steers AND @routes:
$def-conv [definition, converse] 1 ~@want | ~@steers OR ~@routes:
It solves @want 0.763, which is P(steers AND routes) exactly (0.899
times 0.849; network reference, 2026-09-08; under D36 it read 0.755,
within the ridge’s 1/w of the product), and the lint is silent. The
same two lines spelled as forward plus backward ($def-back 1 @steers
AND @routes | @want) carry the same logic, but they form a directed
cycle, so the lint reports W2, and the backward line concludes into the
roots, so it draws W25 on each; and since 2026-09-08 the checkers reject
the spelling with E11: a line concluding more than one statement is
unimplemented under the shipped solve until block normalization lands,
and the message hands you the split, one line per statement at strength
1 (the retired D36 solve read it 0.752, the same as the pair). Nothing in
the semantics prefers that spelling; the converse pair is the one to
write. One-way is
not a definition: a single forward line leaves the term free above the
members (the reference fills the interval: the balancing prior under
D36, the network under D161), so a consumer of the term reads more than
the members deliver (in the toy battery’s t9c, @ab 0.896 against
the members’ product 0.792 under the network reference, measured
2026-09-08; 0.828 against 0.749 under D36 as pinned 2026-08-22). Both
directions, or none.
A pure terminology node is the shape to avoid: @asi-def [ASI means an
AI far beyond every human at every cognitive task] 1: conjoined into a
premise (T7) draws an edge into every junction that uses it and says
nothing a gloss would not. Under the retired uniform reference it also
had a price: the p = 1 statement solved 0.988 under the ridge, so every
junction it joined paid about 1/400 (@dies 0.849 with the conjunct
against 0.854 without, T7b). Under the network reference (2026-09-08)
that price is gone: the statement holds at 1.000 and @dies reads 0.859
with the conjunct and without it. The rule stands on the clutter alone.
Terminology belongs in a gloss or a glossary surface; only definitional
claims become lines, and a definition is a p = 1 line, never a p = 1
statement.
5. From source text to map
This chapter is the workflow that produced the flagship map, distilled from the re-authoring passes logged in AUTHORING_NOTES.md (2026-07-16 through 2026-07-24). It assumes you are extracting an argument from a source (a book, an essay, a debate); mapping your own argument works the same way with yourself as the source.
5.1 Pass 1: skeleton
Extract structure only. Statements, evidences, refinement nesting, labels, glosses, citations; no numbers (unstrengthed lines are legal and compile-inert).
Work in three sub-passes, in this order, because the two extraction directions fail differently: pure top-down invents structure (hub statements the source never asserted, which then solve near-tautologous), while pure bottom-up extracts each section faithfully and never authors the connective tissue: a buried spine (W21) and shared grounds double-counted across clusters. The 2026-08-03 He-essay extraction hit both in one day.
- Spine first, top-down: transcribed, never invented. If the source
draws its own overview (a section-2 diagram, an abstract’s roadmap, a
title conditional), transcribe it: the top tier, the declared foci
(
focus:), the region list and prefix scheme, recorded as the manifest comment (chapter 8). Where the source asserts its own structure (“this is not a conjunction”, “either route suffices”), that assertion is quotable content and belongs to this sub-pass, not to your judgment. Genre flip: a debate or interview asserts no overview up front; there, extract bottom-up first and write the spine once the meta-shape emerges (usually in the wrap-up). Never fake a spine the source did not assert: an invented spine is the top-down failure mode wearing a checklist. - Regions bottom-up, in source order. Local links under local
conclusions, each step citing its sentence; this is where source
fidelity lives. The decisions to make deliberately:
- Statement granularity: what gets to be a claim. A statement label must be a proposition. If a source item is not a premise-to-conclusion step (a meta-principle, a design artifact, pure framing), keep it as a folded gloss or one node with a comment, not as a conjunct in the inference chain.
- Linked vs convergent:
ANDonly where the step genuinely needs all conjuncts (test: does the inference fail if this conjunct alone is false?). Independent routes are separate evidence lines. If a comment says “independent paths” and the factor saysAND, one of them is wrong. - What each objection targets. For every objection ask: which
inference does this grant, and which does it deny? An objection
to an inference is an undercut conditioning on that
$id; an objection to a claim is a rebuttal concluding~@id. Mis-typing this is the most common structural error in first passes. - Objections travel with their answers. A response cites its
objection, never the reverse, so a map loses rebuttals more easily
than it loses attacks: skim extraction, later trimming, and source
prose itself (objections are stated loudly, answers quietly) all
bias toward attacks left standing unanswered, and the solve then
prices an unanswered attack at its full authored strength. Measured
on the settled-question benchmark (the H. pylori map,
experiments/solver-prototypes/GROUND_TRUTH_PROBE-2026-08-13.md): deleting map lines at random flipped the known-true conclusion in a majority of orders, and the flips were driven by responses dying before their objections. So when you record an objection, hunt for the source’s answer with the same diligence you gave the objection, and when you must cut for scope, cut the objection and response as a pair rather than the response alone. The answered-attack audit (argmap-query, reference inmvp/README.md) lists every attack and who answers it. - Nest each cluster’s internal traffic (grounds, caveats, objection pairs) under its target.
- Reconcile: where the real work is. Shared grounds are discovered in sub-pass 2, not planned in 1; promote each to the home the burial test picks, which is the nearest container covering every consumer, not blindly the document top (a ground consumed only inside one case block homes at that block’s top rank). Merge or partition lines that turn out to share grounds (chapter 7’s overlap repairs). Set the tier per chapter 8’s spine test: the source’s disclosure order is the guide (the He map’s depth 0 is its abstract, depth 1 its section-2 overview, depth 2+ its detail sections), with cross-tier edges kept visible as top-level evidences or coarse hulls. Then run the lint, and the nest-audit readout for fold candidates you missed.
5.2 Pass 2: numbers, blind, by rubric
The numbers should reflect the source’s confidence, not your own and not what makes the map solve nicely. The discipline that keeps this honest (pre-registered for the flagship map as D39):
- Fix a verbal-to-probability rubric before assigning anything: a table from the source’s confidence language to values. The flagship rubrics (AUTHORING_NOTES 2026-07-19) map, for example, categorical repeated assertions to 0.93, flat unhedged entailments to 0.9, “by default” claims to 0.85, “could well” to 0.7; grants of an opposing point take the conceder’s register. Keep two tables, one for statement registers and one for inference-step language (the flagship’s statement classes vs R-STEP), even if the values happen to coincide. Where the source is silent, a role default applies: a value your rubric assigns to a structural role rather than to any phrase (a default for an unhedged asserted step, one for an objection the source raises to deflect, and so on); define these in the rubric itself, because you will need them. Convention: the rubric lives in a comment block immediately after the frontmatter.
- Assign all values before the first solve, and do not move them afterwards. If the solve surprises you, the finding is about the argument (or the rubric), and it should be recorded, not tuned away.
- Mark provenance. Every rubric-derived number carries
?. A bare number is reserved for values the source states as a credence or probability (“ten to twenty-five percent extinction odds”). A stated frequency or rate (“below one fatal accident per twenty million flight hours”) is not a credence: keep it in the gloss and derive the statement’s value from the assertion’s register as usual. A stated number on a derived statement goes in a trailing# check:comment, never a pin (rule 4.3.2). - The composition rule when hedges stack: the outermost hedge governs. When two readings are defensible, author the weaker and log both. The same rule covers a source that asserts one proposition in two places at two registers: author the weaker register, note the stronger in the gloss or log.
5.3 Pass 3: review
The seven correction classes that were actually needed, in review- checklist form (every one was discovered as a correction, not foreseen; AUTHORING_NOTES 2026-07-19):
- Strength provenance. Does every strength trace to source
confidence language through the rubric? Is
?on everything rubric-derived? - Undercut target typing. Per objection: which inference does this grant, and which does it deny? (Policy objections wearing implication-undercut shape were the flagship’s most instructive mis-typing.)
- Overlap double-counting. For every same-polarity convergent pair: merge, factor out the shared span, reroute an instance-of to the shared ground, or leave independent and say so in a comment (silence is indistinguishable from an unaudited pair).
- Connectivity. Every non-headline statement should feed some evidence. Dangling sub-conclusions are usually missed links. A norm the source argues for should not stand unargued in the map.
- Nesting. First passes come out flat. Fold clusters under their local conclusion; keep cross-cluster shared ground at top level.
- Coverage. Do a full-source pass before calling the map faithful; record the source’s scope in the frontmatter. Summarizing from memory under-extracts.
- Mechanical smoke. Run the lint, the real parser (open the file in the editor), and a headless solve (chapter 10) before calling it done.
Two additions from later passes:
- Defeat presupposition. For every response/rebuttal: which epistemic state does its ground presuppose? If a response only works while X is undemonstrated, conjoin the statement that says so (a guard), so the defeat lapses in worlds where X is demonstrated.
- Multi-voice overlap. When two speakers concur non-diametrically, do not give them independent convergent lines. Full concurrence is one line at the weaker register; a subset relation is the shared span plus a residual increment elicited conditional on it; an instance supports the shared ground, not the downstream conclusion; genuinely disjoint mechanisms stay independent with a comment saying so. A joint line quotes both speakers saying the sentence: if its gloss has to argue that one of them concurs, he has not, and a grant of the other speaker’s point is joint only when the granted sentence is also that speaker’s own words on the map (a grant of a program the other never states as that sentence is the granter’s line). Read a quote to its full stop before any part of it carries a line; a sentence cut at a comma carried a joint line for one afternoon on the debate map before its second half put it back in one voice (AUTHORING_NOTES 2026-09-17 and 2026-09-18).
- A spoken refusal to price. On a multi-voice map, when a speaker
says on the record that they will not put a number on a claim, put
their words on that statement as a
>quote line and list the claim under their key in the frontmatter’sdeclines:block (section 3.9’s display hints, D164). Their view then shows the refusal where a number would stand. Without the quote on the statement the entry is rejected.
5.4 Quoting and citation discipline
Keep verbatim spans to at most one sentence, roughly 25 words, normally one per node; never alter a quote silently; never reproduce a self-contained creative unit (a parable, a poem) whole; retell and compress instead. Whole-map verbatim budget from any single work: low hundreds of words, and for a short source proportionally less (a tenth of the source is far too much regardless of the absolute count).
Where the span goes is the 3.12 test. A quote that supports the claim
belongs on its own > line with a [^ref] locator; a verbatim phrase
that is a grammatical part of the gloss sentence stays in the gloss in
plain double quotes, optionally echoed by a > line carrying the full
sentence. The budget above counts both. (The older convention of marking
in-gloss quotes ~"…" is retired, and W20 flags any survivors.)
6. Labels and glosses
Statement labels and evidence labels do different jobs.
- A statement label is the claim itself, a proposition, and may be a full sentence. Statement labels are not length-linted.
- An evidence label states the step. It names its subject and
says, in plain words and one clause, what the premises give the
conclusion: “a tiny target and imprecise training make alignment
hard”. (Sharpened 2026-09-25, Felix’s ruling on the flagship’s
@align-hardbox; AUTHORING_NOTES 2026-09-25, the label pass. It read “the headline warrant, in a phrase” until then.) The reason is the layout: a line sits where the eye lands first and its premises in a stack the reader skips, so the label is read before its premises, and cold. Four consequences:- Never the warrant alone as a fragment (“the hardness route”, “concedes the abstraction, denies the rescue”), and never a bare “it” whose noun is on another card (10.4 item 12).
- No figure of speech unless it is the source’s own image and the gloss unpacks it (“a tiny target, aimed at with blunt tools” was neither).
- Never a restatement of a premise’s label (“what can be built, gets built” under the premise “What physics permits, someone eventually builds”): a reader who skipped the premise gains nothing, and one who read it reads it twice.
- Every strengthed line gets one. An unlabelled line renders the first sentence of its gloss, which was written as depth, not as a headline, or its humanized id when the gloss is empty; the older advice to leave an obvious deductive connector unlabelled is withdrawn for strengthed lines.
Objection lines speak in the objector’s voice (“the worry is just chatbots, so no ASI is in prospect”), response lines in the answer’s, and a battery that marks its voice keeps the marker (“the hope: …”). Evidence labels crop at about 56 characters in the graph (lint W5) and must fit their plate at the node’s size tier (TRANSLATION_NOTES L14, pinned by
label-crop.test.ts), so distill: where a plain clause cannot fit, keep the subject and the verb and let the gloss carry the rest. - The gloss is the depth tier: the full reasoning, qualifications, asides, source voice, quotes. Glosses are never length-linted.
The three-job test for evidence gloss text, from the corpus survey that
motivated evidence labels (AUTHORING_NOTES 2026-06-12): gloss content is
either (a) a role tag (“undercut of …”), which is derivable from
topology and should be deleted; (b) the warrant, which belongs in the
label; or (c) format-meta commentary, which belongs in a # comment.
What remains after the test is the genuine depth tier. Job order matters
even inside a gloss: put the substantive point first, because displays
crop from the end.
Two further conventions from the accessibility passes:
- Plain-first, technical-nested: write the gloss in plain language; move a technical restatement to a folded continuation line beginning “technical reading: …”.
- Rubric provenance is not reader content. Elicitation citations
(“R-STEP S2: …”) go in a trailing
#comment on the node line, not in the gloss. Reader-valuable quotes and footnote refs stay in the gloss. - In multi-speaker maps, prefix evidence labels with a speaker tag (“A:”, “L:”, “AL:”); IDs are invisible at graph junctions, so the label carries attribution.
7. Structural idioms
The patterns below carry most of the flagship map
(experiments/llm-extraction/iabied-comprehensive-en.argmap; line
numbers are as of 2026-07-25 and may drift, so each entry also names the
anchor to search for). Excerpts are trimmed; open the real file for the
full context. All excerpts are fragments, not standalone files.
The same catalog appears in compressed form in the skill’s
## Structural idioms section, numbered to match these subsections
(7.1 = idiom 1, and so on). For a pattern in a complete small file
rather than as a fragment, examples/README.md maps each example to the
idioms it demonstrates.
7.1 The objection/response triple
The workhorse. An objection statement (the hope or doubt), an objection evidence concluding against the target, and a response undercutting the objection evidence:
# fragment - not standalone (flagship ~line 628, anchor "@c11-readthoughts")
@c11-readthoughts [We'll read the AI's thoughts and catch bad plans] 0.15?:
$c11-readthoughts-obj 0.2? ~@wont-solve-in-time | @c11-readthoughts:
$c11-readthoughts-resp [punishing visible bad thoughts hides them] 0.85? @wont-solve-in-time | @steering-finds-subversion AND $c11-readthoughts-obj: training against legible bad thoughts selects for concealment, not for good ones
FAQ-shaped sources map one-to-one onto rows of these. The -obj/-resp
ID suffixes are a mnemonic convention, not syntax. Each line takes its
own kind (4.1): the objection line’s is usually the hope’s or the
critic’s (hope, testimony), the response’s the book’s own reason
(mechanism, deductive, or analogy where it answers with a case),
and the undercut takes the kind of its own step, never the objection’s.
A gloss-less objection line in a battery has no words of its own to
classify; it takes the kind of the objection’s ground, read off the
premise statement’s words (the rule the flagship’s classification
settled on, 2026-09-08).
7.2 The undercut ladder, including second-order undercuts
Rebuttal, undercut, response, and undercut-of-undercut are one schema applied repeatedly (flagship ~line 283, anchor “$uc-counting”):
# fragment - not standalone
$uc-counting [this argument form fails in ML contexts] 0.3? ~@fragile | @nn-generalize AND @sgd-bias AND $fragile-count:
$uc-uc-counting [the ML rescue may not transfer to alignment] 0.7? @fragile | @gen-not-values AND $uc-counting:
The second line reinstates @fragile exactly to the extent the first
line’s rescue fails.
7.3 Linked and convergent, side by side
One hub with both shapes (flagship ~line 269, anchor “$fragile-ev”):
# fragment - not standalone
$fragile-ev 0.9? @fragile | @orth AND @contingent AND @fragility:
$fragile-count 0.7? @fragile | @counting: lottery-ticket prior over goal-space
The first is a linked three-conjunct rule (all needed); the second is an independent convergent sibling on the same conclusion. The test for linked: neither conjunct alone suffices. The flagship’s cleanest statement of that test (anchor “$spread-ev”): “neither end alone shows disagreement; together they are the spread”.
7.4 Convergent siblings instead of a false AND
When a source presents overdetermined routes (“any one of these suffices”), write separate evidences, not one conjunction (flagship ~line 170, anchor “$adv-speed-ev”):
# fragment - not standalone
$adv-speed-ev [speed alone breaks the human range] 0.9? @ai-advantages | @adv-speed:
$adv-copy-ev [copyability alone breaks the human range] 0.9? @ai-advantages | @adv-copy:
$adv-selfimp-ev [self-improvement alone breaks the human range] 0.75? @ai-advantages | @adv-selfimp:
The flagship originally had these as a four-way AND; the repair note in the file records why that was wrong (the book is explicit that no single advantage is necessary).
7.5 Coarse summary plus refinement
The whole book’s case is one coarse line whose refinement holds everything (flagship ~line 261, anchor “$link “):
# fragment - not standalone
$link [the book's claim as one coarse implication] 0.93? @everyone-dies | @if-built: unfolds below into the full case
The coarse strength on a refined line is not a solver input (the refinement replaces it); it is the evidence-side check, displayed against what the refinement delivers. Recommended practice: author the coarse strength as your holistic judgment of the whole implication before trusting the steps; the comparison is a free audit.
What the refinement delivers, and why it sits lower than you expect. The folded line shows the composition of the block under it: the steps composed the way the block wires them. Steps the argument all requires multiply. So a refinement of four required steps at 0.9 each shows 0.656 on the fold, and one of two steps at 0.9 shows 0.810 (both measured 2026-09-12 on a chain standing on its own). That is the price of writing the steps down, and the first time you see it the number looks like a bug. It is the arithmetic of a chain: every step is one more thing that has to hold. A conclusion the rest of the map argues against reads lower still, because the fold is read on the solved joint like every other number.
There is a second factor, and it pushes the same way. The folded reading asks whether the steps are in force and whether what they need holds, given the premises the coarse line itself carries. A refinement usually reaches for grounds the coarse line does not carry (the interior roots its steps stand on), and it pays for each of them. A fold whose refinement is written on firm ground and short chains sits close to its coarse number; a fold four steps deep over five borrowed grounds sits far below it. Both are honest readings of what you wrote.
The rule, stated positively: pick the coarse strength at or below the
weakest step you are about to write under it. A chain of required steps
can never come out above its weakest step, whatever you believe about how
the steps hang together: each step is one more condition on the same
event, so the conjunction is at most as likely as the least likely of
them. This is arithmetic and there is no way to author around it. If the
summary you want to write is firmer than any step under it, put the
summary where a summary belongs: a # check: on the conclusion, which
records your holistic judgment and lets the solve be audited against it,
with the line’s own strength set at or below the weakest step.
Three honest repairs when a coarse number sits above its weakest step,
and the source decides between them. The step may be under-elicited, in
which case go back to the passage and read its register again. The
summary may be your holistic judgment rather than a claim about this
chain, in which case it is a # check:. Or, most often, the source gives
several grounds and the map wired them as one chain: write them as
convergent lines into the conclusion (7.4), and the fold saturates
instead of multiplying.
A family of parallel objections that all fail for one reason takes that
reason as a premise. The same reading has a consequence for the other
direction. When a refinement holds k objections that the source answers
with one move, the objections OR upward: each one is another way for the
block to deliver, so the folded number climbs above what you meant the
family to carry. The repair is 7.9’s shared latent conjunct: name the one
reason they all fail as a statement, and conjoin it into every member of
the family, so the members stand or fall together the way the source says
they do. The flagship’s $hope-harmless is the worked example and the
one left undone: four hopes that a superintelligence’s goals never touch
us (a digital realm, a cosmos elsewhere, no evolved greed, boredom), and
the book answers all four with one fact, that material resources serve
almost any goal. The coarse line carries that fact as its premise; the
four objections under it do not, so the fold reads 0.268 against an
authored 0.07 (measured 2026-09-24 under the one draw; 0.313 on
2026-09-12, before it). Judge the repair by the conclusion’s
own value rather than by the fold: a conjunct that repeats the coarse
line’s own premise cannot show up in the folded number, because the fold
is already read given that premise.
The check. argmap-query fold-audit lists every strengthed refined
line in a file with its authored strength, the weakest required step of
its refinement, the premises the refinement introduces that the coarse
line does not carry, and a flag where the authored strength is above that
weakest step:
node mvp/packages/parser/bin/argmap-query.mjs my-map.argmap fold-audit
It is advice and never a diagnostic: a summary written in the author’s own holistic register is a legitimate thing to write down, and the audit only tells you that it is what you wrote.
A layout-driven special case is the coarse hull (see the spine test, 8): when the fine conjunction mixes one cross-region premise with hubs that belong inside the region’s fold, condition the coarse line on just the cross-region premise. The refinement holds the fine line and the local clusters; the collapsed view keeps the cross-region edge.
7.6 The complementary partition (“even if”)
Two routes that would overlap are made disjoint by conjoining the negation of the other route (flagship ~line 638, anchor “$mwb-time”):
# fragment - not standalone
$mwb-hard [the hardness route] 0.9? @misaligned-when-built | @align-hard:
$mwb-time [the timing route, in the solvable worlds] 0.9? @misaligned-when-built | @wont-solve-in-time AND ~@align-hard:
The ~@align-hard conjunct is the source’s own “even if alignment were
solvable” made explicit; without it the two routes double-count.
7.7 The balancing evidence
A conditional is vacuous outside its slab, so a map whose every evidence
on @c conditions on @a says nothing about the ~@a worlds. If the
source asserts the converse, name it (flagship ~line 254, anchor
“$no-doom-otherwise”):
# fragment - not standalone
$no-doom-otherwise [if nobody builds it, it kills nobody] 0.9? ~@everyone-dies | ~@if-built: the book's own converse of the title conditional
This is the converse of 4.2: it speaks to the worlds where the premise fails, so it moves the conclusion there and leaves the premise alone, and it keeps a contested base-rate claim on the map, right of a given bar, instead of hiding it in a prior.
Close the set on a spine line (D168, 2026-09-27). A conjunctive
line into a claim the map leads with has one failure branch per
premise, and the solve fills every branch nothing speaks to at one
half, which then shows as part of the headline. So give each premise
its inverse where one holds: by the meaning of the two claims (#
kind: formal, 0.99, as for the pillar inverses), or in the source’s
own words (the sentence’s rubric row and kind). Where neither holds,
invent nothing: the fill stays, and the band beside the number shows
its share. The flagship’s $link-fine conjoins @if-built,
@misaligned-when-built and @mis-ext; $no-doom-otherwise covered
the first branch, and the other two now read:
# fragment - not standalone
$aligned-no-extinction [an ASI aimed at what we want does not kill everyone] 0.9? ~@everyone-dies | ~@misaligned-when-built: # S2, chapter 12's flat "would"; kind: deductive
$harmless-no-extinction [if a misaligned ASI is not lethal, not everyone dies] 0.99? ~@everyone-dies | @misaligned-when-built AND ~@mis-ext: # by meaning; kind: formal
Keep a by-meaning line strictly by meaning: the second line’s
@misaligned-when-built conjunct keeps it off the aligned worlds, where
~@mis-ext says nothing about what an aligned ASI does (without the
conjunct the misalignment drag reads 0.032 in place of 0.044, the
extra drop being no argument of the book’s). Find the shape with the
drag a reader is likeliest to make: before these lines, zeroing
@misaligned-when-built left the title claim at 0.427, nearly all of
it fill, and about 9 of the 0.708 points at rest were fill too; after
them the title reads 0.617 at rest and 0.044 under that drag
(measured 2026-09-27, AUTHORING_NOTES of that date; 0.603 at rest and
still 0.044 under the drag since the objection rewiring of 2026-09-28,
D170).
7.8 Conditioning on an inference (the designed W1)
A policy conclusion that hangs on an implication, not on a fact (flagship ~line 770, anchor “$shutdown-ev”):
# fragment - not standalone
$shutdown-ev [if built means everyone dies, no one may build] 0.93? @shutdown | $link:
Conditioning on @everyone-dies instead would be subtly wrong: doom
that were unconditional would justify no ban. This is the rare positive
evidence-as-premise; the lint fires W1 by design, and the file says so
in a comment.
7.9 The shared latent conjunct (one doubt, many hopes)
When k objections are expressions of one underlying doubt, name the
doubt as a statement and conjoin it into every member; elicit its prior
once, family-holistically (flagship ~line 435 onward, anchor
“$hope-care”; the shared conjunct is ~@no-right-care):
# fragment - not standalone
$obj-cheap 0.7? ~@not-preserved | @cheap-keep AND ~@no-right-care: a sliver of care plus a negligible bill would get paid
$resp-cheap [it would need a reason to pay ours] 0.9? @not-preserved | @needs-motive AND $obj-cheap:
Prefer as the latent the statement the support side already denies, so attack and support quantify over the same worlds. This was the only structure (of six probed) that stayed stable as hopes were added.
7.10 The epistemic-fact reification (norms and credal thresholds)
“A risk no one can bound justifies a ban” is a threshold argument over a credence, which the solver cannot represent directly (facts about credences are not world facts). The pattern: reify the evidence-state as a first-order statement, state the norm as its own statement, and let a near-deductive step combine them (flagship ~line 793, anchor “@risk-unbounded”):
# fragment - not standalone
@risk-unbounded [No one can currently bound the extinction risk below the actionable threshold]: a fact about what has been demonstrated, not about anyone's opinion
@no-gamble [Running a risk no one can bound below the threshold is impermissible] 0.93?: the norm, stated where it can be attacked
$shutdown-fine [unbounded risk + the norm license "don't build"] 0.9? @shutdown | @risk-unbounded AND @no-gamble:
$bounded-escape [a demonstrated bound would dissolve the case] 0.9? ~@shutdown | ~@risk-unbounded:
Note $bounded-escape: the author naming the condition under which
their own conclusion lapses. A self-declared off-ramp is both honest and
persuasive. The wrong shapes (conditioning on the chance itself, an OR
over world-types) are documented as fixture
examples/edge-cases/e15-reified-chance.argmap.
7.11 Rebuttal guards (which epistemic state does the defeat presuppose?)
Seven flagship responses only work while the risk is undemonstrated, so
each carries the guard conjunct @risk-unbounded AND ... (flagship
~line 688, comment anchor “risk-conditional rebuttal guards”). Structure
only, no new numbers: in worlds where the risk is demonstrated bounded,
the defeats lapse and the objections revive. Ask this of every response
you author (checklist item 8).
7.12 Exclusive alternatives and authored abduction
Rival explanations that cannot both hold are a premise-less constraint
factor plus an authored abductive step
(examples/09-exclusive-causes.argmap, the pattern catalog):
# fragment - not standalone
$who [an eaten cake means one of the two ate it] 1.0 @alice OR @bob | @cake: the abductive step, stated as a contestable rule
$notboth [they would not both have eaten it] 1.0 ~@alice OR ~@bob: premise-less unconditional constraint factor
The lesson recorded there: abduction is authored, not free. Pinning the effect gives the causes no diagnostic lift by itself; “it must have been one of them” is a premise, and making it a visible, attackable node is the point.
7.13 The direct-assertion pattern (spoken and debate sources)
A flat spoken assertion with no stated grounds becomes an attributed
premise-less evidence (experiments/llm-extraction/debate-tang-shapira.argmap):
# fragment - not standalone
$a-blur [A: attention is a blur of what causes what] 0.9? @opaque: the quadratic self-attention transformer "literally is a blur of what causes what" [^t005008]
Sixteen of these carried the debate map. Related: a refusal to give a number (“P(doom) is not assignable”) needs no special syntax; leave the marginal blank and, if the refusal is itself argued, map that argument as an undercut cluster against assignability. How several such lines on one statement combine (pool into one number, or stack, each adding its own weight) is decided (D166 item 4: a line with no premise is a line whose premise is every case, so same-side lines stack where nothing opposes them) and applied after the release: until then the solve pools them, and the lint’s I4 note on every premise-less strengthed line says so. A line that reports one source names the source as a premise instead.
7.14 The parable at zero depth
Narrative belongs in folded gloss continuation lines, not in nodes (flagship ~line 165, anchor “a parable”): a whole illustrative story attaches under one statement, costs no graph structure, and folds away. Use this for the source’s most persuasive prose, which is usually exactly the material that does not decompose into premises.
8. Writing large maps
The flagship map holds roughly 200 statements and 270 evidences at depth 5. The disciplines that made that possible (AUTHORING_NOTES 2026-06-18 onward):
- Width, not depth. Every new objection cluster is a sibling under
its target, never a deeper chain. When a sub-debate wants an eighth
level, promote the deep node to a shared top-level node instead.
Node count can triple while max depth stays flat. Width is for
distinct considerations (2026-09-25): a sibling line that restates a
consideration the map already carries is a second dock of it, and two
same-side lines stack, so it is counted twice. Give its passage a
second
>quote on the existing line instead (10.4 item 18). - A manifest comment block at the top, past about 150 nodes: the
coarse spine drawn in ASCII, every shared node listed with its home
region and consumers, and the region-prefix scheme stated
(
@c5-trade,$c5-trade-obj,$c5-trade-resp). Build in dependency order; lint after every region. - Reuse shared grounds aggressively, and annotate each reuse site with a comment naming the home region; otherwise later editors bury cross-references. A handful of high-traffic shared nodes is what keeps maximal coverage finite.
- Home shared nodes above the clusters that use them. An evidence folded inside a refinement contributes no edges while folded, so the spine and shared grounds must not be buried inside clusters. A node declared inside a cluster whose every edge leaves it is a “stranded node” (validator W6); re-home it with its consumer.
- Folding is a source-structure decision. Where you nest is where readers’ fold boundaries are. Author clusters as refinements under their target; keep shared material outside.
Nesting discipline
Nesting looks like an art but is mostly a mechanical review pass. Real
arguments cluster on their own; first drafts nevertheless come out flat
(checklist item 5), typically with the clusters already visible as #
section-heading comments. Treat that as the diagnostic: section
headings are nesting debt. A divider comment organizes the text file;
only indentation organizes the reader’s view. If you felt the need for a
# ---- divider, the argument just told you where a fold boundary is.
Three tests turn the debt into structure:
- Fold-unit test. Would a reader want to collapse this sub-debate to one line? Then give it a wrapper evidence whose refinement holds the cluster (7.9’s hope-battery shape), and author the wrapper’s coarse strength as your holistic judgment of the cluster’s net force; the refinement-vs-coarse comparison then audits you for free (7.5). Smaller version: a statement’s grounds and their evidences nest under the statement.
- Burial test. Anything referenced from outside the cluster moves up out of it. A shared ground homed inside one cluster still works, but it renders as a cross-reference burial and, in the worst case, a stranded node (W6). Home shared nodes above every cluster that uses them.
-
Spine test. The collapsed view must already show the argument’s shape: a folded evidence contributes no edges, so a buried spine disappears from it, and the validator says so (W21, with the linking evidences to lift named in the message). But do not over-correct into lifting every sub-conclusion to top level: that trades a wall of disconnected cards for a crowded one (the He extraction did both in one day: first zero top-level evidences, then fifty top-level cards). Author the top tier deliberately, and keep it coarse: the headline, the sinks, the major route hubs, and the shared grounds the burial test already forces up, roughly 15-30 cards on a large map; every statement hub consumed only within its own region lives one fold down, inside that region. If the source draws its own overview map (a section-2 diagram, an abstract’s roadmap), the flat view should be that overview.
The mechanics rest on a folding asymmetry: a statement block folds to nothing, an evidence refinement folds to a visible coarse line with its edges intact (7.5). So an edge between two top-level statements must never sink into a statement block. When all its premises are top-level, the evidence simply stays top-level. When it mixes one cross-region premise with region-local hubs (the shape that otherwise forces those hubs to stay top-level and crowds the tier), write it as a coarse hull: a coarse line conditioning on just the cross-region premise, with the fine conjunction and the local clusters in its refinement (
$takeover-ev @takeover-doom | @unaligned-asiin the He map, fine five-way conjunction one level down). The spine edge stays visible collapsed, the detail unfolds in place, and the solve runs on the fine line while the hull’s own number becomes the composition of its steps (the folded readout, 10.3), whose gap to the authored coarse strength is then an audit of the summary, never an error to tune away.Quick checks:
grep -c '^\$'returning zero on a multi-statement map means no spine at all (W21 fires); a top rank past ~40 cards means the tier is set too fine (nothing fires; this one is on you). The ~40 bound assumes one argument. If the map is an atlas, several blocks with::groups organizing them, read the bound per group: a table of contents is wide on purpose, andnest-auditsays so rather than calling it a crowd. One or two free-standing exhibit nodes beside a visible spine are fine (W21 stays silent then); fifteen are not a view, they are a deck of unshuffled cards.Run
argmap-query nest-auditonce the skeleton stands, and again after any restructuring pass: it counts the top tier for you, names the boxes whose opened view is a wide and deep wall, and lists the statements whose support cone is ready to fold, each with the edges that block the fold (reference inmvp/README.md). It counts a cone by what still stands at the statement’s own tier, so once you fold part of a cone under one of its own members the suggestion goes quiet instead of repeating itself. It is advice, not a check: nothing it prints is a diagnostic, and declining a fold it proposes is a normal outcome. Two siblings run the same way.argmap-query group-auditnames boxes where three or more objection families stand without a::grouparound them and prints the group header to paste, and it flags a declared group whose lines tell no one story (a possible catch-all).argmap-query cook-auditnames structure that would steer the solve rather than report it: objections buried two refinement levels below the claim’s supports, and near-duplicate supports stacked on one claim. Both are advice in the same sense, and the call stays yours.
Sometimes all three tests fail and the heading is still real. That
happens when the section is a topic, not a fold unit: two arguments
that share a file but not a single premise, or a shelf of background
facts. Nesting them under a wrapper evidence would be a lie: there is
no inference there to summarize. That is what a declared group is for
(3.11): write ::id [Label] and indent them under it. The heading stops
being debt and becomes a checked, drawn box, and the validator will tell
you if the topics you claim are separate have quietly grown a shared
premise.
A useful smell figure: the flagship map holds 431 nodes at depth 5. A hundred-node map at depth 1 is under-nested even if every individual line is well-formed; its reader meets a wall of top-level nodes and the fold control does nothing.
9. Limitations
What the format and semantics currently cannot express, with the standing workarounds. None of these block parsing or display; they bound what a solve can mean.
- Scope conditionals (FORMAT_DESIGN Q8). “Aligned in the current
regime, degrades at superhuman scale” has no first-class form.
Marginals capture partial truth, not the conditioning scope; nesting
is a partial workaround whose limits fixture
e06documents. The intended direction (D14) is partitioning statements into substatements; undesigned. Until then, statement granularity is the author’s burden. - Undercut fan-out. An undercut names one target. Class-level methodological objections (“this is all unfalsifiable”) attack a family of inferences and end up structurally under-stated as one representative undercut. Mitigations: give the family a shared gate premise and rebut that once; or make the objection a shared Tier-1 ground feeding several undercuts.
- No statement re-opening (D33). You cannot declare a statement and attach its refinement later in the file; refinement is physical indentation. Workaround: declare nodes at their refinement site and forward-reference them (IDs are document-global).
- Statement-level provenance.
?marks numbers as estimated, but who asserts a claim has no in-format home beyond footnotes and ID prefixes. In multi-source maps this makes scope policing (“does this node belong to this map’s source?”) a manual discipline. - Binary statements only (S8). Categorical or continuous claims must enter through threshold-gate statements (“X exceeds T”).
- Credal links are second-order (S12). Nothing computes “if the
probability of X exceeds t then Y”; the epistemic-fact reification
(7.10) plus a
# gate:comment audit is the pattern. - Facts, norms, and future scenarios mix by convention only (S13). The flagship keeps policy conclusions as “should” statements and has no “will X happen” node whose truth would feed back onto its own antecedents. If you add scenario nodes, index them explicitly or the map becomes self-referential.
- Independence is assumed (P2) and dependence must be authored (4.4). There is no correlation annotation.
- Solver cost grows with treewidth (P5). The corpus solves in seconds at treewidth about 7 to 9; a much more entangled map may not. Width-not-depth authoring also keeps treewidth down.
- Comment-layer slots are conventions.
# check:and# gate:are invisible to tools other than the solver readouts, and nothing validates them structurally. - Cross-map ID reuse is unchecked. Reusing an ID across maps is
string coincidence; verify the propositions match before treating
them as the same claim (a debate’s “prepare an off-button” is weaker
than the book’s
@shutdown). - Authoring cost is real (RISKS §2). Mapping is slower than prose. The mitigations that exist today are LLM extraction with human steering, and the rubric discipline that keeps the numbers honest (RISKS §4); neither removes the labor, they redistribute it toward review.
10. Checking your map
Three mechanical gates, in order: lint, parse, solve.
10.1 The lint
python3 tools/argmap-lint.py path/to/your.argmap
Zero errors is mandatory. The codes (full table in tools/README.md):
| Code | Meaning | Author action |
|---|---|---|
| E1 | duplicate ID (one namespace across @/$) |
rename |
| E2 | dangling reference | fix the ID |
| E3 | ~$id |
rewrite as an undercut (3.6) |
| E4 | probability outside [0,1] | fix |
| E5 | v0.3 pair without argmap-version: 0.3 |
declare the version |
| E6 | malformed pair (0.9/, /0.2) |
write both members |
| E7 | ::id in an expression (3.11) |
a group takes no part in inference; reference a member |
| E8 | a probability on a :: line (3.11) |
groups have no credence slot; delete the number |
| E9 | > outside an annotation block (3.12) |
move the quote under its node, before that node’s first child |
| E10 | > with no quote text |
write the quote or delete the line |
| W1 | evidence-in-premise, not undercut-shaped | usually a polarity slip; legitimate only for deliberate conditioning-on-an-inference (7.8), then say so in a comment |
| W2 | directed cycle | usually fine (mutual rebuttal); check it is not a zero-negation support cycle |
| W3 | prose line resembling a node | you lost a sigil; fix it |
| W4 | footnote used/defined mismatch | fix |
| W5 | evidence label past ~56 chars | distill the warrant; depth to the gloss |
| W7 | pair sums > 1 | declared two-sided conflict or infeasible residual; confirm intended |
| W8 | pair 0/0 |
drop it |
| W9 | pair entangled with undercut shape | check what the opposed side actually asserts |
| W11 | authored 0 strength | you probably mean an unstrengthed line |
| W15 | quote line with no [^locator] (3.12) |
add the locator; provenance is the point |
| W16 | leftover [^ inside quote text (3.12) |
only a trailing ref is the locator; fix the stray or doubled one |
| W17 | ` # ` inside quote text (3.12) | quote lines have no trailing comment; move the note to a #[…] line |
| W18 | quotes before the end of the gloss prose (3.12) | reorder: gloss first, then quotes |
| W19 | > under a declared version below 0.3 |
declare argmap-version: 0.3 |
| W20 | retired ~"…" still in a gloss (3.12) |
migrate it: > line, plain marks, or the echo pattern |
| W21 | buried spine: top-level statements unconnected in the flat projection (8) | lift the linking evidences it names to the top tier |
| W22 | zero-node file | the parse went wrong; read the counts, fix the file |
| W23 | a bare point and a # check: on one head line (4.3, 4.7) |
drop one, or write the unargued remainder as a pair beside the check |
| W24 | an @id/$id token in prose that resolves to nothing |
fix the typo, or drop the sigil if it is not a pointer |
| W25 | a bare point beside mapped support (4.7) | derive: move the number to a # check:, or write the remainder as a pair |
| W26 | malformed # check: token (3.7) |
a point in [0, 1] or lo..hi with lo <= hi |
| I1 | stats; isolated statements | connect or delete isolates |
| I2 | block inventory (multi-block files) | read it; confirm the split you intended |
| I3 | unstrengthed-line inventory (4.7) | the lines that compile inert; commit or drop each before calling the map done |
| I4 | a premise-less strengthed line: the reading note (7.13) | the solve pools it with the other numbers on its conclusion for now; the stacking reading is decided (D166 item 4) and lands in a later release; a line reporting one source names the source as a premise |
| I5 | parallel leaves: two or more sibling lines with one conclusion and one premise set (settled by D166) | they are same-side shares of one population, independent where nothing opposes them; one argument written twice merges |
| I6 | coinciding pairs from different premise sets into one statement (settled by D166) | where their premises hold together they are one draw and the strongest share on each side holds (4.4), so agreeing lines read as either one alone |
Two caveats: the stranded-node check (W6) lives only in the TypeScript
validator (visible in the editor), not in this lint; and on
expression-valued conclusions (@a OR @b left of the given bar) the
lint’s undercut-shape family (W1/W9 and same-slab W7) deliberately stays
single-ref, so near-misses there are the editor validator’s job (see
tools/README.md, update 2026-07-25).
Warnings are advisory and some are load markers on purpose: the flagship map ships with two deliberate W1s. The discipline is not “zero warnings”; it is “every warning has an explanation you could put in a comment”.
10.2 The parser
Open the file in the editor (or run the parser test suite if you work in
the repo). The editor shows diagnostics inline, including the
validator-only warnings (W6 stranded node, W10 version gate). Without
the editor (a standalone tool bundle), a successful solve_map.py run
doubles as the parse gate: it loads the file through the real parser.
10.3 The solve
cd experiments/solver-prototypes
python3 solve_map.py path/to/your.argmap --top 10
Read the four sections: statement gaps (authored or check value vs solved), evidence tensions (authored strength vs achieved), spectator gaps (a folded line’s authored strength vs the conditional the whole network delivers through its refinement, a bench readout) and composition gaps (the same authored strength vs the composition of the refinement’s own steps, which is the number the editor shows on the folded line). Then query the nodes you care about:
python3 solve_map.py path/to/your.argmap @headline '$main-step'
(The forced interval is a bench instrument of the retired uniform
reference: add --reference d36 --band. It needs the optional
band_probe.py next to solve_map.py, and
--influence needs influence_probe.py; the plain readout needs
neither.)
Interpreting what you see:
- A large statement gap: the mapped argument does not deliver the authored or checked belief. Revise structure or strengths if the argument is misstated; add named missing evidence if real support is unmapped; otherwise keep the badge, it is a finding.
- A large evidence tension: the map holds the line below the strength its author wrote (under D161 only that direction tints; a line carried above its strength is the fill); look for an overlooked conflict with neighboring lines.
- A composition gap: the refinement’s own steps, with the claims they rest on that nothing inside argues, compose to something different from the coarse strength its author wrote on the folded line (“steps outrun summaries”, or the reverse at the spine). That composition is the folded line’s readout in the editor (since 2026-09-08, D161 item 9), marked by a dotted track and, since 2026-09-12, tinted like any other gap (D38 as amended); decide which side is wrong, both states occur in practice. The spectator gap beside it is a bench readout only: the network’s delivered conditional also reads the conclusion’s other parents, so it is not the line’s own number and not a verdict on either side while SOLVER_SEMANTICS S21 is open.
- Numbers never move to make badges disappear (5.2.2). Structure moves, named evidence is added, or the badge stays and means something.
The solver needs numpy, scipy, and node. If it is unavailable, lint plus parse still validate everything structural.
For translated maps, python3 tools/translation-parity.py BASE TR
verifies the translation touches only free-text spans.
10.4 The pitfalls checklist (for an adversarial auditor)
Work this list against a finished map as if you were paid to find the
double count. Each item says what to look for, how to see it (the lint
code or argmap-query audit that catches it, or the reading that does
when no tool can), and the fix. The tools are
python3 tools/argmap-lint.py FILE and
node mvp/packages/parser/bin/argmap-query.mjs FILE <audit>; every
audit is advisory and prints the shape, and the call stays yours. The
rationale for each item is 4.7.
- A pin beside mapped support. A statement with an authored point
and at least one strengthed line concluding into it; the point
counts that line’s consideration twice (T1). See it: W25. Fix:
derive it (the number moves to
# check:at its register’s interval), or write the unargued remainder as a pair. - A pin and a check on one head line. The point pins the statement to itself while the check silently disagrees; the solve reports a tiny gap that reads as convergence. See it: W23. Fix: drop one; on an interior statement the check stays and the point goes, or becomes a residual pair if the text licenses one.
- A pinned root homed in a box it does not feed. A pinned root
declared inside a refinement box or
::group, consumed only outside it, wired to nothing inside it: shared ground parked where the map holds its likeliest antecedents. See it:isolate-audit(each row names the container and the outside consumers). Fix: derive it from the container if the inference test passes, re-home it with a consumer, or say in the gloss why it stays. - A restatement wired as inference. A line whose premise and
conclusion are the same proposition or episode at two docks (the
@psychosis/@c13ws-retrainshape): one consideration counted twice through a line that carries no inference. See it:restate-auditlists lines whose premise and conclusion cite only the same locator; the reading is “does the passage make this step, or say the same thing twice?” Fix: delete the line and cross-reference the docks in a comment, or merge them into one statement. - Two lines sharing a premise into one conclusion. Same-polarity
lines into one statement whose premise sets overlap, combined as if
independent. See it:
shared-cause, one row per shared statement. Fix: either the overlap is the deliberate world-layer device (the shared conjunct, 7.9) and a comment says so, or factor the shared span out, or merge the lines. - A statement AND-ed with its own derivative. A junction
@x AND @ywhere some line derives@yfrom@x(the flagship’s$wst-race): the consideration enters the junction twice, and no static audit sees it. See it: the reading. For every AND, trace each premise’s ancestry two hops up (argmap-query neighbors ID --hops 2); a premise that appears in another premise’s ancestry is the hit. The covariance probe (experiments/solver-prototypes/covariance_probe.py) finds it at scale as premises that co-move under one root’s ablation. Fix: drop the duplicate premise, or re-elicit the line conditional on the derivative alone. - A floor elicited from the total. A derived statement carrying a
residual pair whose floor sits at, or near, the register’s point:
the pin again with extra steps (T3’s hazard; the spoke’s variant C).
See it: the reading. A pair beside a check is lint-silent by design,
so compare the two: a floor that would close the badge on its own is
the total, and a floor with no passage behind it in the gloss is the
total by default. Fix: empty the residual unless the text names a
second unwired ground (floor at that ground’s register, passage in
the gloss) or a remainder (“to name a few”:
0.2?/0?). - A posterior elicited as direct evidence. A root’s pair read off a register that already discounts what the author believes downstream, so the solve applies modus tollens a second time. See it: the reading; the signature is an author hedging a root because of a conclusion the map also derives from it. Fix: there is no exact correction. Keep the table’s widths, say in the trailing comment that the register is a posterior, and present the interval and the band rather than the midpoint.
- A definition as a p = 1 statement. A terminology node
(
@asi-def [ASI means ...] 1) conjoined into premises: it solves 0.988 under the ridge and taxes every junction it joins by 1/400 (T7). See it: grep the marginals for ` 1:and1?:`; the lint is silent (0/1 is a smell without a code). Fix: terminology goes in a gloss or a glossary surface; a definition that does inferential work becomes p = 1 evidence lines, spelled as the converse pair (T6). - A malformed check interval.
# check:holding a token that is neither a point in [0, 1] norlo..hiwith lo <= hi (reversed bounds, a bound past 1, a single dot, a dangling..). The check is display-only, so nothing fails downstream; the badge simply disappears while the comment reads as if it counted. See it: W26. Fix: write the point or the interval. - An unstrengthed line left as structure. A line with no strength compiles inert: the relation is drawn and nothing is asserted, so a map that “maps” an argument through it says nothing about it. See it: I3, the inventory. Fix: commit a strength by rubric, or delete the line if the relation was never meant to carry.
- A label whose pronoun has no antecedent. “It must hold across
the gap between weak before and lethal after” (the flagship’s
@before-after-gapbefore 2026-08-23) reads cleanly in source order and as nothing in the outline, the graph card or a search hit. See it: the reading. Read every statement label cold and out of order (sort the outline, or read theargmap-query rootslisting); any label that needs the line above to say what “it” is fails. Fix: put the subject in the label (“Alignment must hold across the gap …”). - A bundled label. A label carrying two propositions (“X and Y”),
or a claim fused with what follows from it (“X, so dismiss Y”): two
variables in one slot, and a reader cannot tell which half a line
argues for. See it: the reading; grep labels for ` and
,so,therefore`, and ask of each hit whether the two halves could be true separately. Fix: one proposition per statement, or derive the fused conclusion from its halves with one line per ground. - A refusal written as a number. The source says “nobody knows”
and the map carries 0.5, or a confident point, or a narrow pair,
with no gloss. See it: the reading; a pair wider than 0.30 with no
gloss sentence, or a 0.5 with none, is the signature. Fix:
0.1?/0.1?plus the gloss sentence naming the passage, or a blank marginal where the refusal is the author’s own. - A pair whose shape contradicts its register. A categorical the
source leaves no room against, written two-sided; an objection the
source denies, written
0/pwith no floor; a hedged claim written as a flat pair. See it: the reading, against the table in 4.7.4 and the class token in the trailing comment (a pin with no class token is unaudited). Fix: the table’s row, or a per-node pair from the text with the rationale beside it. - A restatement under two labels, and the wrong door. A line whose
premise and conclusion describe one event under two labels
(
@asi-soon“ASI arrives this century” feeding@if-built“anyone builds an ASI” on the flagship until 2026-09-25): the line carries no inference, its gloss tends to carry the source’s real reason (there, the chapter 12 race) that its premises do not name, and a what-if that zeroes the premise leaves the conclusion at the reference fill, which a reader sees as a broken sum (50%). Its companion is a claim whose only consumer is that line (@int-power, chapter 1’s power, reached the title only through the build), so a whole chapter’s case enters the argument through the wrong door. See it:restate-auditonly where the docks share a locator; otherwise the reading, which the drag makes cheap: for each strengthed line into a hub, zero its premise (solve_map.py --override) and read the hub; a landing near 0.5 is the shape.consumerson each top-level statement finds the single-consumer case;inverse-audit --alllists the hubs a check prices and no inverse line does. Fix: give each label its own proposition (feasibility here, the build there), put the source’s reason in the premises, wire the orphan where the source uses it, and write the analytic inverse where one holds by the meaning of the two claims (~@if-built | ~@asi-soon, deductive, no passage needed). Rule: AUTHORING_NOTES 2026-09-25. - An evidence label that is not the step. A strengthed line with
no label, or one whose label is a warrant fragment, a figure the
gloss does not unpack, or its premise’s label said again. The reader
meets the line before its premises (section 6 item 2), so none of
these tells them what the step is about or where it lands. See it:
the reading. List every strengthed line’s head sorted by label, so
each is read without its neighbours:
grep -oE '^\s*\$[A-Za-z0-9_-]+ (\[[^]]*\] )?[0-9.]+\??' FILE | sed -E 's/^\s+//' | sort -t'[' -k2(the rows with no bracket sort first: those lines have no label; on the flagship, 363 rows and 82 of them bare before the 2026-09-25 pass, 0 after). Ask of each: does it name its subject, and does it say which conclusion the premises give? Fix: one plain clause, premise to conclusion (“our intelligence remade the planet, so it is powerful”), in the objector’s voice on an objection line. Parallel lines into one claim need different words:cook-auditreads a label overlap of 0.75 or more as a duplicate support. Rule: AUTHORING_NOTES 2026-09-25, the label pass. - One consideration at two docks. Two strengthed lines into one
claim that are one argument read twice: one motive under two labels
(
$ext-coreand$ext-incidentalinto@mis-exton the flagship until 2026-09-25), a rule beside its instance ($soon-framebeside$built-ev), one answer the source gives in two paragraphs (the hired hands and the robot bodies into@wed-lose), or a premise that restates the conclusion ($fragile-count). Same-side lines are independent draws that stack where nothing opposes them (D161, D166), so the second dock counts the consideration twice and the claim reads above what the source’s one argument delivers; at rest a saturated claim hides it, so a drag does not find it either. See it: the reading, andshared-causeis its trigger: every row it prints (lines into one claim sharing a premise, outright or through a refinement) is a pair to read against the source, asking whether the two passages are one answer.cook-auditcatches only near-identical labels, and a pair with disjoint premises (a rule and its instance, one answer in two paragraphs) shows on no instrument, so read the lines of every multi-line conclusion side by side against their passages. Fix, by the rule one consideration, one dock: a second passage on the same consideration becomes a second>quote on the same line, never a second line; where a rule feeds an instance, the rule becomes one statement with its own box and one line takes it; where two passages are one answer, the lines merge into one (conjoin the premises, or OR them where either suffices: one line, one coin); where a line sits at the wrong door, re-aim it. Never nest the fix inside an evidence. Rule: AUTHORING_NOTES 2026-09-25, the cleanup.
Run the mechanical half first (lint, then the five audits: isolate-audit,
restate-audit, shared-cause, cook-audit, pinned-roots, and
inverse-audit --all for the hubs a drag would expose), then the
reading half over the statements pinned-roots lists, since those are
the rows where items 3, 7, 8, 14 and 15 live, the drag over the hubs
for item 16, the sorted label listing for item 17, and a side-by-side
reading of every multi-line conclusion for item 18, starting from the
rows shared-cause prints.
Appendix A: a complete worked example
The file below is complete and lints clean as shown (zero errors, zero
warnings). It exercises: convergent
routes, a linked conjunction, a refinement with an evidence-side check,
an undercut, a reinstating undercut-of-the-undercut, a rebuttal, check
credences, ? discipline, and a footnote.
---
argmap-version: 0.2
title: "Protected bike lanes and cyclist safety"
description: "AUTHORING_TUTORIAL.md Appendix A: worked example."
date: 2026-07-25
---
# Headline first (convention). Derived statement: no authored marginal,
# a check credence instead (residual authoring rule).
@lanes-safer [Protected lanes reduce cyclist injuries per trip]: the headline claim # check: 0.8
# Route 1: observational. The coarse line refines into the per-trip
# reading; its 0.7? is the evidence-side check against the refinement.
@study-drop [Injury rates fell after protected-lane installation] 0.9?: city-level before/after counts [^lusk]
$obs-route [before/after data carries the claim] 0.7? @lanes-safer | @study-drop: unfolds into the per-trip reading below
@exposure-ok [The drop is not explained by reduced cycling] 0.8?: ridership rose over the same period, so per-trip risk fell
$obs-fine [per-trip injuries fell while ridership rose] 0.8? @lanes-safer | @study-drop AND @exposure-ok: linked - both facts are needed for the per-trip reading
# Route 2: mechanism. Convergent sibling of $obs-route (independent
# routes, so separate lines, not an AND). Independence audited: the
# mechanism does not rest on the before/after data.
@separation [Physical separation removes the main collision type] 0.9?: most serious urban cycling injuries involve motor vehicles
$mech-route [the design removes the dominant injury mechanism] 0.75? @lanes-safer | @separation:
# The objection: grants the data, denies the inference from it
# (an undercut of $obs-route, not a rebuttal of the claim).
@confound [Cities add lanes where cycling is already safest] 0.5?: selection: lanes go where streets are calmest
$uc-obs [selection could explain the before/after drop] 0.6? ~@lanes-safer | @confound AND $obs-route:
# The response: an undercut of the undercut (reinstatement). The
# selection story predicts no drop at quasi-random sites.
@natural-exp [Some installations were sited quasi-randomly] 0.7?: construction-driven and court-ordered sitings
$resp-uc [quasi-random sites show the same drop] 0.8? @lanes-safer | @natural-exp AND $uc-obs:
# A rebuttal (attacks the claim itself, so no $-conjunct).
@risk-comp [Riders take more risks when they feel protected] 0.4?: the risk-compensation hypothesis
$rebut [risk compensation could offset the design gain] 0.3? ~@lanes-safer | @risk-comp:
[^lusk]: Lusk et al., "Risk of injury for bicycling on cycle tracks versus in the street," Injury Prevention 17, 2011.
What to notice:
@lanes-saferis derived, so it carries# check: 0.8and no authored marginal.- Frontier roots (
@study-drop,@separation,@confound, …) keep authored values, all?-marked as estimates. $uc-obsconditions on$obs-route(undercut);$resp-ucconditions on$uc-obs(reinstatement);$rebutconditions on neither (rebuttal).- The refinement under
$obs-routemakes its 0.7? a displayed check against what$obs-finedelivers, and$uc-obsre-aims onto the refinement’s delivery line when unfolded.
The actual solver readout for this file (solve_map.py at its default,
the network reference G that ships since D161, with the one draw of MATH
§3.8; measured 2026-09-23), abridged:
appendix-a.argmap: 7+5 vars, width=6 | 0.0s, conv=True, reference=g, cell-draw=chain
largest statement gaps (authored/check -> solved):
@lanes-safer 0.80 -> 0.848 |d|=0.048
@exposure-ok 0.80 -> 0.799 |d|=0.001
...
largest spectator gaps (authored coarse ~> delivered by refinement):
$obs-route:delivered-by-refinement p=0.70 ~> q=0.856 |gap|=0.156
largest composition gaps (authored coarse ~> in force by composition, the folded readout):
$obs-route:in-force(composition) p=0.70 ~> q=0.549 |gap|=0.151
badges: 0 statements more than 0.10 outside the check interval
Reading it: the mapped argument delivers a little more than the check
credence 0.8 on the headline (solved 0.848, inside the 0.10 badge
threshold), so the map carries the stated belief; whether the author’s
0.8 was too cautious or the map is missing a qualifier is the question
the badge would ask if the gap grew. The rebuttal is what holds the
headline there (4.4): where risk compensation holds, $rebut stays in
force in 0.25 of those cases, near its authored 0.3, and the headline
reads 0.737 there. Under the rule before 2026-09-23, which weighed the
routes and the rebuttal as independent evidence, the routes outvoted it
(in force in 0.10 of those cases, the headline 0.836 there and 0.890
overall; both measured 2026-09-23 on the solved joint). The sub-0.01 gaps on the roots are the solve meeting each point
to its resolution, not tension. The composition row is the number the
editor shows on the folded $obs-route (10.3): the refinement’s own
steps, with the claim they rest on that nothing inside argues
(@exposure-ok, 0.8) and the selection undercut where its answer fails,
compose to 0.549, well under the authored 0.7. The coarse strength
promises more than its own steps deliver, which is the “decide which
side is wrong” case of 10.3: either the 0.7 is too generous, or the
exposure claim deserves an argument of its own. The spectator row above
it is a bench readout (solve_map.py prints it; the editor’s reader
surface does not): the conditional the whole network delivers through
the refinement, 0.856 here, which also reads the headline’s other route
($mech-route) and is therefore not the line’s own number
(SOLVER_SEMANTICS S21, open). Under the retired uniform reference
(--reference d36, the readout this appendix quoted until 2026-09-08)
the same file reads @lanes-safer 0.818, the composition 0.520 and the
delivered conditional 0.837 (re-measured 2026-09-23; the composition
figures this paragraph quoted until then, 0.690 and 0.668, predate the
composition’s counting of a route’s unargued interior premises, which
landed on 2026-09-08).
Appendix B: cheat sheet
@id [label] p?: gloss statement; p optional, ? = estimated
$id [label] s? CONCL | PREM: gloss evidence; s optional; | optional
$id s ~@x OR ~@y: gloss premise-less constraint factor
~@id negation (never ~$id: E3)
AND / OR linked / convergent; parens to mix
::id [label]: gloss declared group; no credence, never in an expr
indented node lines under @/$: refinement (replaces parent unfolded)
under ::: membership in the group
indented prose folds into the gloss above
> verbatim text [^locator] quote line; one line, no trailing comment
# comment full-line or trailing
#[key: ...] annotation comment (per-file free in parity)
# check: p or # check: lo..hi display-only credence (derived stmts); badge = distance to the interval
# kind: word on an evidence line: formal, deductive, mechanism, empirical (n=<int>), testimony, analogy, hope
# gate: q($e) >= t => @c threshold audit (comment layer)
[^ref] ... [^ref]: source footnote citation
---: argmap-version: 0.3 required for slash pairs (s+/s-, p+/p-) and > lines
Number rules: elicit as “assume the premises; how likely is the
conclusion?”; ? on rubric-derived values, bare only for source-stated
numbers; derived statements get checks, not pins; no authored 0/1; fix
arguments, not numbers, after the first solve.
Counts (D161, 2026-09-08): a claim’s firmness is its width (0.7/0.1
8 flips, 0.85?/0.05? 18, a point 200, the cap); a line’s is its
# kind: (formal hard, deductive 1000, mechanism 64, empirical
and testimony 16 or a larger stated n=, analogy and hope 4, no
key 16). The tint lights only outside what was written (below a
strength, outside an interval, off a point, past 0.01); inside is the
fill, uncoloured. What-if: your number replaces the author’s on that
claim as a point at the cap; on a root the map follows forward, on a
conclusion the author’s case retreats where it is softest (free
premises, hopes and analogies, judgments, flat assertions, mechanisms,
in that order); “reached” on the adjustment row is the refused residual (W4-E renamed it from “held at” on 2026-09-08, so “held at” is the count’s phrase alone; “met at” became “reached” on 2026-09-27).
Undercut schema: $u q ~C | grounds AND $target. Ask: which inference
does this objection grant, and which does it deny?
Review checklist, one line each: provenance traced; undercut targets typed; overlaps merged/factored/partitioned or declared; no dangling sub-conclusions; clusters nested, shared grounds top-level; full-source coverage pass; lint + parse + solve smoke; defeat presuppositions guarded; multi-voice overlaps deduplicated.
Check: python3 tools/argmap-lint.py FILE, then
cd experiments/solver-prototypes && python3 solve_map.py FILE --top 10.