ArgMap authoring tutorial

Audience: a person who wants to read or write .argmap argument maps. Format version: v0.2 (syntax frozen 2026-07-09; the v0.3 slash-pair extension is noted where relevant). Semantics: ratified D36, amended by D161 (the 2026-09-08 semantics campaign: the network reference as the default solve, a count on every number, the kind key on lines, the set rule for the tint, the what-if as revision).

Agent-facing companion: the argmap-author skill (.claude/skills/argmap-author/SKILL.md) compresses this tutorial into the working loop; agents load it via the Skill tool, humans can read it as the cheat-sheet-plus. Since 2026-07-27 the skill is meant to be sufficient on its own for authoring (it carries the idiom catalog, the label/gloss and nesting discipline, the limitations and the lint codes in compressed form), so the division of labour is: skill = the pattern, this tutorial = the rationale, the history, the reading chapter, and the worked example in Appendix A. A third tier, examples/README.md, indexes the example corpus by idiom for when you want to see a pattern in a whole file rather than as a fragment.

This tutorial distills the project’s design docs and the accumulated authoring experience into one document. It never overrides them: on any point of doubt, FORMAT_DESIGN.md (syntax), GRAMMAR_DRAFT.md (grammar), SOLVER_SEMANTICS.md (semantics), and GLOSSARY.md (terminology) are authoritative, and DECISIONS.md records why things are the way they are. MATH.md is the readable account of the mathematics the semantics rests on (what the numbers mean formally, what has been proved about them, and what is still open), and is the right next stop after chapter 4. AUTHORING_NOTES.md is the dated log this tutorial condenses; new learnings continue to land there first.

Snippet convention: every .argmap block in this tutorial is either a complete file that passes tools/argmap-lint.py as shown, marked (complete, lintable), or an illustrative fragment whose first line is # fragment - not standalone.

Contents:

  1. What ArgMap is
  2. Reading argmaps (self-contained; you can stop after this chapter)
  3. The format
  4. What the numbers mean
  5. From source text to map
  6. Labels and glosses
  7. Structural idioms
  8. Writing large maps
  9. Limitations
  10. Checking your map
  11. Appendix A: a complete worked example. Appendix B: cheat sheet.

1. What ArgMap is

ArgMap is a plain-text format plus an editor and viewer for making complex arguments explorable. Instead of reading a linear essay, a reader navigates the argument as a graph: the main claim and its support are visible at a glance, and every reasoning step can be unfolded to the depth the reader wants. The motivating use case is AI safety argumentation, where the arguments are long, branching, and full of objections that attack specific inference steps rather than conclusions. The flagship content is a comprehensive map of the book “If Anyone Builds It, Everyone Dies” (IABIED), deployed at p1graph.org.

An .argmap file describes a bipartite factor graph with two node kinds:

  1. Statements (written with the @ sigil) are variables: propositions that can be true or false, optionally annotated with the author’s credence that they hold.
  2. Evidences (written with the $ sigil) are factors: reasoning steps that connect statements, optionally annotated with a reliability. An evidence line is one evidence written on one line of the file; readers see it as a reasoning link.

Roles such as premise, lemma, and conclusion are never declared; they are derived from the graph topology (a statement nothing points into is a premise, one nothing points out of is a conclusion). Attacks are not a separate primitive either: an objection is an ordinary evidence whose conclusion is a negated statement, and an attack on an inference (an undercut) is an evidence that references the attacked evidence itself. This uniformity is the core design idea: two node kinds and one reference mechanism express support, opposition, rebuttal, undercut, and refinement.

The text file is the single source of truth. The editor renders it as an outline and a graph, but everything those views show is derived from the text, and everything you author happens in the text.

The v0.2 syntax is frozen (DECISIONS.md D25 to D33). Anything this tutorial shows is stable; future syntax changes arrive as versioned format changes (the first is the v0.3 slash pair, gated by an explicit argmap-version: 0.3 frontmatter declaration).

2. Reading argmaps

This chapter is for readers: people who explore existing maps in the viewer or query them from the command line. It does not assume or require anything from the authoring chapters.

2.1 The viewer

The editor/viewer at p1graph.org has three synchronized panes: the text (the .argmap source), the outline (a collapsible tree of the same content), and the graph. Reader and focus views present single nodes and their neighborhoods in a more article-like form. The graph starts collapsed: boxes with a fold control contain refinements, finer subgraphs that replace a summary reasoning step when unfolded. Folding follows the source structure, so what unfolds together is an authorial decision, not a layout heuristic.

Conventions worth knowing when reading:

  1. @ nodes are claims; $ nodes are reasoning steps between claims.
  2. An evidence pointing at a claim supports it; an evidence pointing at a negated claim opposes it. An evidence that takes another evidence as an input attacks (or conditions on) that inference itself, not its conclusion.
  3. A number on a claim is the author’s asserted probability that it holds. A number on a reasoning step is its reliability: roughly, how likely the step is to actually carry when its inputs hold. A trailing ? marks a number as estimated rather than deliberately asserted; in the file only, since 2026-09-08: the popup’s numeral no longer carries the ?, its sentence rides the hover text and the hollow caret on the gauge, and the popup shows instead a firmness beside every number, a count in coin flips on a claim and the step’s kind word on a reasoning step (2.2).
  4. In the flagship map, numbers derive from the book authors’ own confidence language through a fixed rubric (DECISIONS.md D39), so disagreements the display surfaces are audits of the source’s coherence, not the map maker’s opinions.

2.2 Implied values and tension

The viewer can compute what all the authored numbers jointly imply. Under “Show what the map implies” (in the Controls popover; on by default since D40, though an explicitly persisted opt-out still wins), an in-browser solver reads every authored number as a claim and finds the distribution that honours all of them together, as far as they can be honoured at once. Since D161 (2026-09-08) it starts from the map’s own network of reasoning steps rather than from ignorance: each line is read as a row of a probabilistic network, and the authored numbers are soft targets on that network (before D161 the solve was the maximum-entropy distribution honouring the numbers as constraints, each a spring rather than a wall, D36). Each node then shows an authored -> implied readout; the editor calls the computed number the implied value, and the technical documents call the same number the solved value.

Every number has a second axis. A number says how likely; beside it the viewer shows how firmly it is held, in the currency of coin flips (a number held at n flips gives way under pressure as far as a rate estimated from n tosses would). The firmness is never a third thing to author. For a claim it comes from the width its author left open: an interval 0.7/0.1 (the claim somewhere between 0.7 and 0.9) is about 8 flips, a flat point 198 flips (about 200; the chip prints 198), which is the cap every point gets. For a reasoning step it comes from the step’s kind, read off the source’s own words: a mechanism holds 64 flips, a record or a judgment 16, an analogy or a hope 4, the source’s own deductive claim 1000, and a formal step is hard. The popup shows the kind word on a step and the count on a claim; the detail readout says the same in a sentence. Chapter 4 says where the numbers come from (4.1 for steps, 4.7 for claims).

The tint means conflict. A readout colours only where the implied value has left what the author wrote: below a step’s strength, outside an interval, off a point, and by more than 0.01 (the solver’s own resolution at a point, so a point met to 0.005 stays uncoloured). Everything else is the solve filling what the author left open, and it is shown without colour. On the flagship as authored that is the whole picture: all 239 authored intervals are met inside their bounds, 112 of them more than 0.05 above their lower bound (both counts re-measured 2026-09-24 on the shipped solve, solve_map.py at its default over every pair statement; 224 and 113 on 2026-09-08, before the map’s later content and the one draw), and none is tinted. A tint therefore says exactly one thing: the rest of the map could not honour this number, and the size of the tint is how far it fell short. Some statements carry a check credence (a displayed comparison value that does not constrain the solve); the badge comparing it to the implied value has the same meaning, the distance from the implied value to the check’s interval. On the flagship one sits past 0.10 as authored, and it is the headline: @everyone-dies reads 0.603 against 0.75..0.95, 0.147 short, an audit finding about the book; the next widest is its third premise @mis-ext, 0.795 against 0.85..0.95, a hair ahead of @fragile at 0.796 (measured 2026-09-28, after the two objections that deny @mis-ext were aimed at it, D170; 0.617 on 2026-09-27, after the two inverses on the title claim took out the reference fill that had held the headline at 0.708, D168). The headline’s check was re-read on 2026-09-25 from the book’s unconditional sentences, with five other checks, each from the sentence that prices its own quantity; the badge past 0.10 before that, @evo-analogy’s, closed on 2026-09-08 when the chapter’s own case for the analogy was mapped.

Moving a number yourself. In what-if mode a reader drags a claim’s number. The map re-solves with the reader’s number in place of the author’s on that claim, held as firmly as an author’s point (198 flips, the cap). On a claim nothing argues for, the map follows the reader forward and nothing colours. On a claim the author’s argument delivers, the reader’s number is a new claim beside the author’s case, and what bends shows what that case holds least firmly: whatever carries no number at all first, then, other things equal, in the order of firmness, hopes and analogies (4 flips), wide intervals, records and judgments (16), the author’s flat assertions (18 on the flagship), mechanisms (64), points (200), the source’s own deductive claims (1000), and a formal step never. The order is a tendency by count, not a rule by position: only numbers on the path between the reader’s claim and the case’s roots can give, so a mechanism on that path bends before a hope off it (4.6 shows one such case). Where the map cannot honour the reader’s number, the adjustment row says by how much: on the confusions map’s Socrates syllogism a “not mortal” at 0 reads “reached 32%”, the two premises giving to 0.67 each and the formal step holding (4.6, re-measured 2026-09-08 through the shipped solve; the row’s exact wording is of 2026-09-27, when “met at” became “reached”). Section 4.6 walks through both cases with the flagship’s numbers.

2.3 The headless readout

To query a map without a browser, use the CLI readout (from the repo):

cd experiments/solver-prototypes
python3 solve_map.py ../llm-extraction/iabied-comprehensive-en.argmap @shutdown '$link'
python3 solve_map.py ../llm-extraction/iabied-comprehensive-en.argmap --top 10

The first form prints, for each named node, the authored value and the solved value. The second prints the ten largest gaps and tensions in the whole map: the places where authored numbers and computed numbers disagree most. --reference d36 --band adds the forced interval for a named statement (how far the constraints actually pin it, as opposed to where the solver settled inside the allowed range); it is a bench instrument of the retired uniform reference, so it has to be asked for with that reference and its numbers do not describe the shipped solve. --condition if-built=0 --id @everyone-dies answers a what-if by conditioning (the named statement taken as a fact about the world, the author’s numbers updated by Bayes), the bench reading that the viewer’s what-if mode does not use (4.6); --override is the viewer’s reading, revision. Quote $id arguments so the shell does not expand them. The solver needs python3 with numpy and scipy, and node on the PATH.

That is everything a reader needs. To write maps, continue.

3. The format

An .argmap file is plain text, UTF-8, with optional YAML frontmatter, comment lines, node lines, and footnote definitions. Indentation is spaces only; a tab is a parse error.

3.1 Statement lines

@id [short label] p: gloss
  1. @id declares a statement. IDs use [A-Za-z0-9_-], are case sensitive, and share one namespace with evidence IDs: @x and $x cannot coexist. Prefer mnemonic IDs (@risk-unbounded, not @s17).
  2. [short label] is optional. For statements the label is the claim, phrased as a proposition. If no label is given, the gloss serves as the display label.
  3. p is optional: the author’s credence that the statement holds, a probability literal in [0,1]. A trailing ? (as in 0.7?) marks the value as estimated or unelicited rather than deliberately asserted.
  4. Everything after the : is the gloss: one logical line of free text giving depth, sourcing, or qualifications.

A statement is a variable, so its label must be a proposition, something that can be true or false. “Anyone builds it” is a statement; “the question of whether anyone builds it” is not. Conditionality lives in evidences, never in statement labels. Normative propositions (“X should happen”, “doing Y is impermissible”) are legal and ordinary statements; what chapter 9’s fact/norm caution forbids is not norms but future-fact nodes that would feed back onto their own antecedents.

3.2 Evidence lines

$id [label] strength <conclusion-expr> | <premise-expr>: gloss
  1. $id declares an evidence: a reasoning step asserting that its premises bear on its conclusion.
  2. [label] is optional and carries the headline warrant: why the premises support the conclusion, in a phrase (chapter 6).
  3. strength is optional: a bare probability literal, the evidence’s reliability (chapter 4 explains precisely what it means). ? works as on statements.
  4. The | is the given bar and reads “given”: $e @c | @a is evidence about @c given @a. The conclusion side is a full expression, not just a single reference.
  5. The | <premise-expr> part may be omitted entirely. A premise-less evidence is an unconditional constraint factor: it asserts its conclusion expression with the given reliability, unconditionally. Example from the spec: $rivals ~@hyp-fluke OR ~@hyp-filter: rival explanations can't both hold.

3.3 Expressions

Premise and conclusion sides use the same grammar:

  1. References: @id for a statement, $id for an evidence (see 3.6), each optionally negated with ~ (~@id). Negating an evidence reference (~$id) is forbidden: it parses, but the validator rejects it (error E3). To challenge an inference, write an undercut (3.6).
  2. AND joins linked premises: the step needs all of them.
  3. OR joins convergent premises: any one suffices.
  4. Mixing AND and OR requires parentheses: (@a AND @b) OR @c. Unparenthesized mixing is a parse error; there is no silent precedence.
  5. Unicode ∧ ∨ ¬ are accepted as input aliases; the canonical form is ASCII. & is not a connective (the | character is taken by the given bar, so OR cannot be written | either).

Note the graph-level route to convergence: two separate evidence lines with the same conclusion are independent factors, which is usually the right way to say “two independent reasons” (see 4.4 for when it is not, and for what happens to two reasons once their conclusion is known).

3.4 Refinement (nesting)

Indentation always means “belongs to the line above”. What belongs means is read off the parent line’s sigil: under @ and $ it is refinement (this section); under :: it is membership in a declared group (3.10). Everything below is the @/$ case.

An indented block under an evidence is a refinement: a finer-grained subgraph that models the same reasoning step at higher resolution. When a reader unfolds the evidence, the block replaces it; folded, the outer line serves as the coarse summary. One consistent indentation increase per level (two spaces recommended).

# fragment - not standalone
$syllogism @socrates-mortal | @socrates-human AND @humans-mortal: surface form
  @intermediate [Socrates inherits mortality property] 0.99:
  $inherit-1 @intermediate | @socrates-human AND @humans-mortal-property:
  $inherit-2 @socrates-mortal | @intermediate:

A statement may also carry an indented block; that refines the implicit factor asserting the statement’s own marginal, and the statement itself remains.

Refinement changes what the folded line reads: a folded evidence shows the composition of the block under it, so the steps you write below a coarse line decide the number the reader sees on it. 7.5 says how to pick the coarse strength so the two agree.

IDs are document-global: a node declared inside a refinement can be referenced from anywhere, and forward references (using an ID before its declaration) are legal. Where you nest is a real authoring decision, not formatting: nesting determines what folds away together in every view (DECISIONS.md D22).

3.5 Glosses and continuation lines

A gloss is one logical line, but it can be hard-wrapped: any deeper-indented line that does not begin with @, $, or # folds into the gloss of the nearest preceding node line, joined with a space. This is also how longer narrative passages attach to a node without costing graph structure:

# fragment - not standalone
@human-precedent [Human intelligence transformed the planet] 0.9?: Nobels to humans, none to chimps
  a parable: a council of beast-"gods" laugh at the Ape-god's newest
  creature, frail and clawless. The Ape-god says only, quietly, "and yet."

The hazard: if you forget a sigil on a node line, the line silently becomes gloss text of the node above. The lint warns when a prose line looks like a node declaration (W3); take that warning seriously.

The inverse hazard has no warning, and cannot get one. A continuation line that begins with @, $, #, :: or > is read as that construct, not as prose, because line dispatch is a first-character switch, and by the time anything could complain, the parser has built a node and has no idea prose was intended. So never start a continuation line with a sigil character: begin with a word, or rephrase. The exposure is small because quote lines (3.12) cannot wrap and the corpus barely uses continuation lines at all, but when it bites there is no diagnostic: you find it by reading the rendered gloss.

Two lexical restrictions: a gloss cannot contain a # preceded by whitespace (that always starts a trailing comment), and a label cannot contain square brackets.

3.6 Evidences as premises: conditioning and undercuts

An evidence named in another evidence’s premise expression denotes that evidence’s activation (“this inference is in force”), not its conclusion. There are two uses.

The positive use is conditioning on an inference: $policy @act | $link makes a conclusion depend on an implication holding, rather than on a fact. This is rare and legal; the lint flags it (W1) because the same shape is usually a polarity mistake, so when you do it deliberately, say so in a comment.

The common use is the undercut. To attack the inference $E C | P (rather than its conclusion), write:

# fragment - not standalone
$E-undercut ~C | <grounds> AND $E: why the inference fails

The conclusion negates $E’s conclusion; the premises conjoin the grounds with $E itself. Conditioning on $E is exactly what makes this an undercut rather than a rebuttal: if $E is itself disabled or undercut elsewhere, the undercut lapses with it. Dropping the AND $E turns it into a plain rebuttal, which fires regardless. Undercuts of undercuts (reinstatement) are the same schema applied again.

Coming from probabilistic conditional logic

If you know conditionals of the form (psi | phi)[d] from the maximum-entropy literature (Kern-Isberner, Paris, the SPIRIT and MEcore systems), five translation rules keep the solve honest; each was measured on examples/toys/pcl-penguin.argmap (2026-08-31, and again under the shipped solve 2026-09-24; in the app, the Semantics audit map audit-prior-art, group P1).

  1. (psi | phi)[d] with d at or above 0.5 is the line $e d psi | phi.
  2. With d below 0.5 it is the opposed line at 1 - d: $e (1-d) ~psi | phi. A line’s number is the inference’s reliability, so 0.01 @fly | @peng is a nearly worthless reason for flying, and the solve concludes from “birds fly” that penguins fly (0.95 under the what-if).
  3. A subclass exception is an undercut of the general rule on the subclass plus a rebuttal. 0.9 @fly | @bird beside 0.99 ~@fly | @peng is a contradiction here, because “birds fly” in force is a law over penguins too; with the penguin pinned the map is infeasible. Write 0.99 ~@fly | @peng AND $birds-fly (the undercut, which alone leaves flying at even odds) and 0.99 ~@fly | @peng (the rebuttal, which then reads the textbook 0.01).
  4. An independence statement, a conditional at the base rate such as (allergic | treated)[0.10] beside (allergic)[0.10], has no line form. Leave it out: a support line at the base rate pulls its antecedent down, and the lint flags the attempt (W25). If the pull it was meant to block is real, pin the statement it protects.
  5. A fact (a)[d] is the bare point @a d; whether it should be a pair is the two-floor discipline of 4.7 (W23, W25).

The solver never reads your descriptions, so the same three lines with different glosses solve to the same numbers. What a description can do is tell you which shape to write. Three readings of “birds fly, 0.9” and the shape each licenses:

  • The law. “All birds fly; I am 90 percent sure.” Write 0.9 @fly | @bird and, for every exception you know, the undercut plus rebuttal of rule 3. The rule stays a law over the whole class.
  • The frequency. “90 percent of birds fly.” Either scope the rule to the class it is true of, 0.909 @fly | @bird AND ~@peng (the number raised so the mixture over the class comes back to 0.9) with @peng’s base rate written down, or keep the unscoped line and read the exception off the solved map by conditioning (the CLI’s --condition), never by pinning the exception: a pinned penguin beside an unscoped “birds fly” is a contradiction in this reading as much as in the textbook’s.
  • The case. “This bird probably flies.” A point or pair on @fly itself, or a pinned premise; no rule about birds is being asserted.

3.7 Comments and comment-layer conventions

A line beginning with # is a comment; a whitespace-preceded # starts a trailing comment. Comments are preserved by the parser and serializer. Section headings in large maps are full-line comments by convention. Use one when the heading is only for a human reading the source. When you want the heading to be checked and drawn, a named box around those nodes, use a declared group instead (3.11): no tool can see a comment.

Two trailing-comment conventions carry meaning to the solver tooling without being syntax:

  1. # check: p on a statement line records the author’s all-things-considered credence for display against the computed value. It never constrains the solve. Chapter 4 explains when to use it instead of an authored marginal. A check value may carry the ? marker like any other value (# check: 0.9?), and in a source-faithful map it should: there the check is the source’s own stated register for that conclusion (the D39 practice), not the extractor’s belief. A check may also be an interval, # check: 0.85..0.95: the range the register licenses, at the same widths the residual pairs use. The badge a reader sees is then the distance from the computed value to that range, zero inside it; a bare point is the zero-width interval. A token that is neither (reversed bounds, a bound past 1, a single dot) is the lint’s W26.
  2. # gate: q($id) >= 0.10 => @conclusion records a threshold audit: after a solve, if the left side clears, the named conclusion is expected to hold, and the display reports agreement or disagreement.
  3. # kind: <word> on an evidence line records what sort of step the line is, read off the source’s own words: formal, deductive, mechanism, empirical, testimony, analogy or hope. It sets how firmly the solve holds the line under a reader’s what-if (its count, in coin flips), never its strength; 4.1 says when to write which. An empirical line may add its stated sample size (# kind: empirical n=200), and the key may share a comment with other keys, separated by ; (# check: 0.9?; kind: mechanism). The key is content: it is identical in every language of a translated set (TRANSLATION_NOTES L16, checked by translation-parity.py). On a refined coarse line the key is for the reader; the solve reads the leaves’ keys, because the refinement replaces the coarse line (3.4, 7.5).

All three are conventions, not grammar; tools other than the solver readouts will treat them as ordinary comments.

3.8 Citations

Attach citations as Markdown-style footnotes: [^ref] in a gloss, and a definition line anywhere at top level:

# fragment - not standalone
@p1 [Capabilities advance rapidly] 0.9: doubling times keep shrinking [^epoch2025]
[^epoch2025]: Epoch AI, "Trends in Machine Learning," 2025.

Footnote text is free text: author, title, venue, year, plain URL. Viewers autolink URLs; there is no inline link syntax. The lint checks that every used footnote is defined and every defined footnote is used (W4).

3.9 Frontmatter and file layout

Optional YAML frontmatter between --- fences carries metadata: title, author, date, description, source, scope (what part of the source the map claims to cover, checklist item 6), and argmap-version (declare 0.3 if the file uses slash pairs or quote lines). Unknown keys are preserved, which makes frontmatter the extension point for provenance notes.

One key is tooling-visible: focus: [id, id] (D57) declares the map’s focus nodes, the statements influence readouts measure deltas on. Omit it and the tooling derives them from topology (statements concluded by top-level evidence and premised by none). Declare it only when topology misreads your intent: the known case is a goal guard, a world-layer conjunct like ... AND @shutdown on strategy-advice lines, which makes the goal premise-referenced without arguing from it. The list is complete, not additive.

One key is a display hint: fold-links: on (D162) asks the viewer to open your map with its “Fold links” mode on, so a folded branch shows which other branches it touches. Declare it when your map’s branches cross each other and a reader needs to see that at a glance; leave it out otherwise, since the mode is off by default and it does reshape the folded layout. A reader’s own checkbox in the graph controls still wins for their session.

A second display hint belongs to multi-voice maps: declines: (D164) lists, per speaker key, the claims that speaker refused on the record to put a number on, as a: everyone-dies [^t012346], ids without the @. In that speaker’s view the claim then shows a blank gauge with their words where the view’s arithmetic would stand. Use it only for a spoken refusal, never for a claim the speaker merely left unpriced: every entry needs a verbatim > quote line by that speaker on that statement, and the optional [^locator] picks which one the reader sees. The solve ignores the hint, so everything downstream of the claim keeps its value in that view.

Top-level order is free; the graph defines the structure. Convention: put the document’s headline claim first, then work down its support.

3.10 v0.3: two-sided pairs (brief)

Since D52/D53 a file declaring argmap-version: 0.3 may write two-sided values: on an evidence, $e 0.9/0.2 @c | @a adds an opposed floor toward the negated conclusion in the same slab; on a statement, @s 0.8/0.1 bounds P(s) to [0.8, 0.9] instead of pinning a point. No whitespace around the slash; ? binds per member; an omitted second member is 0 and means exactly the v0.2 reading. New maps can ignore pairs until they need to express “this consideration cuts both ways” or an interval-shaped residual; details in FORMAT_DESIGN §3.1/§3.2 and SOLVER_SEMANTICS §1.9.

3.11 v0.3: declared groups (::)

A third sigil declares a group: a named box drawn around nodes.

::timelines [Capability timelines]: when transformative AI arrives
  @agi-soon [Transformative AI within a decade] 0.4:
  @compute-grows 0.9: frontier training compute keeps growing
  $scaling 0.7 @agi-soon | @compute-grows: the trend argument

Its indented block is membership, not refinement (the one place the indentation rule is keyed on the parent’s sigil, 3.4). Otherwise the head reads exactly like a node line: required id, optional [label], optional gloss, and the same continuation-line folding.

Three properties define it:

  1. No credence. There is no probability slot on a :: line, now or later. A number there is an error.
  2. Not referenceable. ::id in any expression is a parse error. A group takes no part in inference: you cannot argue from it or against it.
  3. Transparent. Deleting every :: line changes nothing about the map: same graph, same roles, same solve. A group is display only.

Use one to say “these nodes are one topic”. Before ::, the only way to say that was a # ==== banner comment, which no tool could see, check, or draw.

  1. A kind box is headed by its thesis (Felix’s ruling, 2026-09-09). A group that boxes objections of one kind takes as its label the shared thesis of its members, in the objector’s voice and general enough that every member is an instance of it: ::halt-enforcement [A halt cannot be made to stick]. A reader arrives with a proposition in their head and matches by content; a neutral question head (Can the halt be made to stick?) makes them translate first. Check every member against the head before committing; where one is not an instance, adjust the head, never the member. The gloss opens on the member list, because the folded card’s popup shows it and it is the second discovery channel. The scope, decided the same day: the rule binds a KIND card nested under an objection shelf; a document-level shelf keeps its question or topic head (Is the worry real and near?, Objections to the halt), which reads well as it is and has no single thesis to state.

Groups and blocks. A block is derived: a connected component of the graph, a set of nodes that reach each other. A group is authored. They usually coincide, and the validator checks the relationship: a group equal to one block, or spanning several whole blocks (“two topics under one heading”), is silent. Two shapes warn, because your claim and the graph disagree:

  • a group covering only part of a connected block (W12): edges cross its boundary, so the open box distorts the layout and a fold can only summarize the crossings;
  • one block split across two groups (W13): usually an accidental cross-topic premise silently merged two topics while your headings still assert they are separate. This is the mistake worth catching.

Only document-level groups are checked. A group nested inside another group, or inside a refinement, is subdividing its parent, not claiming a block.

Folding. Every group folds to a card. A closed group (no edge crosses its boundary) folds with nothing to redraw; a non-closed group folds too, and its boundary edges become dashed summary edges bundled by direction and kind, so a battery of twelve objections folds to one card with one bundled attack edge. The card shows the label, a tally of what its hidden edges do to the surrounding claim, and never a credence: the summary vocabulary is the guard here, since the card only reports where hidden traffic goes and takes no part in the argument itself. A group nested inside a refinement starts folded, so opening a box shows its groups as cards first; a document-level group starts open.

When to group. The workhorse case, measured on the flagship map, is the objection battery inside a box: several answered attacks on the box’s claim, each objection paired with its response. Grouping them turns a wall of interleaved nodes into one labeled card next to the claim, and the claim’s supports become findable again. Guidance from that campaign: keep a family’s lines contiguous so the group is a pure wrapper; put a ground inside the group only if nothing outside the family consumes it, and leave shared grounds outside as boundary premises; give the group one dominant story toward its claim (a box that both attacks and supports its claim in equal measure makes an illegible card); label it as a short reader question (“Is the worry real and near?”) rather than a category noun. Leave a box flat when its objections interleave with shared-ground supports, and treat a group created purely to fix layout as provisional: re-judge it in the reader’s default view whenever the viewer changes.

Groups are allowed anywhere: any depth, inside each other, inside refinements. Membership does not suppress the isolated-statement note: a context shelf of standalone facts still reports each one as isolated.

3.12 v0.3: source quote lines (>)

A line beginning > under a node carries a verbatim span from your source, plus the footnote locator it came from:

# fragment - not standalone
@no-honor [Honor is a contingent evolved hack an AI won't carry] 0.9?: honor is
  an evolutionarily contingent shortcut, not a convergent feature of minds
  > a specific weird hack that humanity stumbled into [^supp-ch5]
  > quite skeptical that gradient descent will happen to stumble across the

(the last line is shown truncated only for the page width, see “no wrapping” below.)

The gloss goes back to being a claim a reader can parse cold; the quotes sit under it as its evidence. Before >, a quote could only live inside the gloss, where no tool could see it. That meant no display affordance, and, in a translated map, nothing stopping a paraphrase from being presented as verbatim.

Rules, all short:

  1. The text is verbatim. Never paraphrase it, never silently repair it. If you need to trim, trim at the ends.
  2. Always give a locator. [^ref] at the end of the line, defined at top level like any footnote (3.8). A quote without provenance is almost always an authoring slip, and the lint says so (W15). Locators are as coarse or fine as your source allows: a chapter ([^epub-ch7]), a supplement page ([^supp-ch5-promises]), a transcript timestamp ([^t001734]).
  3. Only a locator at the very end of the line counts. Everything else on the line is verbatim text, including a [^…] in the middle of it (W16 flags that as a probable stray or doubled ref).
  4. No trailing # comment, the one line kind that has none. Source text cannot be reworded to dodge the comment splitter, so a real ` # ` in a quote would be silently truncated; instead the whole line is verbatim and the lint warns if it spots ` # ` inside one (W17). Put per-quote notes on an annotation comment instead (below).
  5. No wrapping. A quote is exactly one line, however long; the editor soft-wraps it.
  6. Placement is positional. A quote attaches to the node above it and must be indented deeper. It has to sit in that node’s annotation block, the span before the node’s first child. A > at top level, or after a child node, is an error (E9), not a re-attachment to some outer node. Order quotes after the gloss prose; interleaving parses, but the lint prefers the canonical order (W18) and the serializer rewrites to it anyway.
  7. Declare argmap-version: 0.3 in a file that uses > (W19).

Which quotes become > lines: the three-way test. Ask: is this the node’s own wording, or support for it?

  1. Supporting quote: a fragment stacked next to the claim as evidence for it. Lift it to a > line. This is most of them.
  2. Load-bearing inline fragment: a verbatim phrase that is a grammatical constituent of the gloss sentence (Kelvin's "infinitely beyond…" fell to DNA). Leave it in the gloss, in plain quotation marks: it is the node’s own phrasing, borrowing the source’s words. When the provenance is worth keeping, add an echo, a > line carrying the full verbatim sentence and its locator, while the gloss keeps its fragment.
  3. The quote is the claim: the gloss is nothing but the quote. Degenerate case of 2: write the gloss in plain marks and echo the verbatim on a > line.

The echo pattern also keeps translations honest: a translated gloss renders the fragment as ordinary quoted prose (claiming nothing about verbatimness), while the > line stays in the source language.

Quotes are never translated. In a multilingual map set the whole > line (sigil, indent, text, locator) is byte-identical across all language versions, and tools/translation-parity.py enforces that. A translated “verbatim” quote is false on its face and destroys the tie back to the source.

Annotation comments (#[…]). Per-quote side data goes on a full-line comment of the form #[key: …] (no space between # and [) on the line above the quote, at the same indent:

# fragment - not standalone
  #[de: schwer, der Schlussfolgerung zu entgehen]
  > hard to avoid the conclusion [^supp-ch5]

To the parser this is an ordinary comment. Two things make the form worth using rather than a plain #: the parity tool treats #[…] lines as free per file (every other comment must match byte-for-byte across translations), and it is the reserved surface for real attributes in a later format version, so today’s convention promotes without a rewrite. Its current tenant is the parked translation of a quote, waiting for a real translation field.

4. What the numbers mean

The numbers are the part of the format most worth getting right and the part where intuition most often misleads. The ratified semantics (DECISIONS.md D36, full treatment in SOLVER_SEMANTICS.md) reduce to a small set of rules an author can hold in their head.

This chapter gives those rules operationally: what to write, and why it behaves as it does. If you want the model underneath them (what the map compiles to, why the solve is a maximum-entropy problem, and which of these rules are theorems rather than conventions), that is MATH.md, published as The mathematics behind ArgMap. Nothing here depends on reading it.

4.1 Evidence strength

Elicit an evidence’s strength by asking: assume the premises hold; how likely is the conclusion? That prompt is the whole elicitation procedure. Formally the number is the unconditional in-force rate of the rule (how often this kind of inference actually carries), a property of the rule itself, independent of whether its premises happen to be true. The two readings coincide numerically by construction, so you can elicit with the conditional prompt and reason with either picture.

Practical consequences:

  1. The strength isolates the inference. Whether the premises are true is carried by the premises’ own numbers, elsewhere in the map. Do not discount a strength because you doubt the premises.
  2. Contraposition is not a rewrite. $e 0.8 @c | @a and $e2 0.8 ~@a | ~@c are different claims; the given bar is directional. When extracting or translating, preserve the direction the source actually asserts.
  3. A strength of 1 is legitimate for deductive steps: the line becomes a pure constraint (“a proved implication has no reliability coordinate”). A strength of 0 is almost never what you want; the unstrengthed line is the exact “structure only” form (the lint suggests this, W11).
  4. An evidence with no strength at all contributes structure to the display but nothing to the solve. This is a deliberate, useful state: sketch the shape of the argument first, commit numbers later.
  5. A near-certain premise costs nothing. Under the retired uniform reference a line’s strength was held on both sides of its premise, so a line conditioned on a near-tautology (for example the OR of four of five partition members) showed a large spurious tension on the nearly empty side, and this item warned against it. The shipped solve holds a line’s rate once, on its own coin: a 0.8 line on a premise pinned 0.98 reads its strength exactly, where the retired reference read the empty side 0.079 off (a two-line toy, solve_map.py at its default and at --reference d36, 2026-09-24). Condition on the premise the source states.

The second axis: the line’s kind. A strength says how likely the step is to carry. Since D161 (2026-09-08) every line also has a count: how firmly the solve holds that strength when something else in the map pushes against it, in coin flips (a strength held at n flips gives way as far as a rate estimated from n tosses would). The count is never elicited as a number. It comes from the line’s kind, what sort of step it is, which a reader can classify from the source’s own words where nobody could read a count off a book. Write it as the comment key # kind: <word> (3.7). The rubric, with the cue words and the count each kind holds, as fixed on 2026-09-08 (deductive raised from 200 to 1000 flips the same day, after the campaign’s skeptic pass; DECISIONS.md D161 carries the current table, this one is dated):

kind the step cues in the source flips
formal logic, definition, arithmetic, a machine-checked derivation, a universal instantiation; checkable independently of the author a proof, a calculation, “all F are G and x is F” hard (or simply write the step at 1)
deductive the source’s own claim that the conclusion follows “by definition”, “necessarily”, “it follows”, “is the same as” 1000
mechanism a causal or structural reason that would operate whenever the premises hold “because”, “the process”, “would tend to”, “any such system” 64
empirical a frequency or a record “historically”, “in every case so far”, a named count of incidents, a study 16, or a larger stated sample size
testimony the authors’ or experts’ stated judgment, offered as such “we think”, “experts”, “most researchers”, “our judgment” 16
analogy an inference carried by a likeness “like”, “as with”, “analogous”, “the way evolution” 4
hope the source’s own labelled hopes and speculative objections “the hope:”, “perhaps”, “one might hope”, “it is conceivable” 4
no key     16

Four rules ride on the table:

  1. Formal against deductive. A formal step is written at strength 1, or at its strength with # kind: formal; either way the solve holds it hard, a constraint no other content bends. A source’s own deductive claim is a different thing: the authors assert that the conclusion follows, and they can be wrong about their own logic, so the line is firmer than any statement a point can be (1000 flips against the point cap’s 198, five times the cap) and softer than a formal step, which never bends. A reader’s what-if therefore bends a deductive line after every premise a point holds, and a formal line never. Measured on the confusions map’s Socrates syllogism (c4: two premises at 0.99, the syllogism at 0.99, the reader’s “not mortal” at 0) through the shipped solve with the syllogism re-keyed deductive (2026-09-08, solve_map.py --override and the compile’s socrates-tier test): at 1000 flips the syllogism gives 0.06 (0.99 to 0.93) where each premise gives 0.30 (0.99 to 0.69) and the reader’s 0 is held at 0.30; at the 200 it first had, the syllogism gave 0.24 (0.99 to 0.75), nearly as much as its premises, which is what the raise fixed. Under the map’s own formal key the step holds at 0.99 and each premise gives 0.32. Said plainly: on the two-premise copy the deductive step does bend. Its coin reads 0.93 while the premises give, which misses the condition the raise was set for (the coin above 0.95 while the premises give) by 0.02; on the single-premise copy (c2a) the same key holds at 0.986. The formal key holds exactly on both, 0.990. The constant stays at 1000 and the 0.02 is on Felix’s list (D161 item 3; the ladder through the compile reads 0.75 / 0.93 / 0.96 / 0.98 / 0.99 at 200 / 1000 / 2000 / 5000 / 20000 flips, pinned in socrates-tier.test.ts).
  2. The boundary sentence. deductive only where the line’s own words claim necessity, definition or elimination; a reason that would fail if the world were arranged otherwise is mechanism, however confidently the source states it. (Two blind passes over the flagship disagreed on 25 lines at exactly this seam before the sentence existed; with it they resolve by rule.)
  3. A stated sample size raises the count and never lowers it. # kind: empirical n=200 holds a survey of two hundred at 200 flips; a line that reports one incident stays at the default 16. The floor is measured: three premise-less one-incident lines at one flip each held he-xrisk’s @containment-fails at 0.75 against the author’s own check of 0.85, and at the default within 0.013 of it (PLACEMENT_PROBE section 23 (e), 2026-09-08).
  4. An undercut or rebuttal takes the kind of its own step, never of the line it attacks: an undercut by mechanism is mechanism, and “technophobia, not an inference” is testimony.

What the count changes, and what it leaves alone. The map as authored barely moves: the flagship’s 348 keyed lines move one statement past 0.05 (@nat-abs 0.14 to 0.19) and light no new badge (2026-09-08; under the one draw, 2026-09-24, still one, @nat-abs by 0.051, and no badge). What the count decides is which line gives first under a reader’s what-if (4.6). The solve reads four counted tiers (1000, 64, 16, 4) and the hard one, so a disagreement between empirical and testimony, or between analogy and hope, moves no number; the word is still worth getting right, because the reader sees it on the popup. The flagship’s 348 lines carry their keys since 2026-09-08: 105 mechanism, 69 deductive, 65 testimony, 45 hope, 34 empirical, 30 analogy and no formal (the book has no machine-checkable step). The map has grown to 361 lines since, every one keyed (2026-09-24: 108 mechanism, 69 deductive, 68 testimony, 48 hope, 36 empirical, 32 analogy). Since 2026-09-26 two of them are formal: the analytic inverses $not-feasible-not-built and $no-win-no-extinction, each true by the meaning of its terms, at 0.99 (recount that day: 106 mechanism, 69 testimony, 66 deductive, 48 hope, 35 empirical, 31 analogy, 2 formal). Since 2026-09-27 a third line is formal, $harmless-no-extinction, one of the two inverses on the title claim (7.7, D168). The rubric’s record, with the two-pass agreement (kappa 0.65) and the third reader on the ties, is ideas/plans/semantics-antecedent-pull-2026-09-06.md section 17.

4.2 The unpriced ground, and when to price it

A line says nothing about its own premise. If the only line in a map is $imp 0.8 @c | @a, the solved P(@a) is 0.500 and P(@c) is 0.700: the network gives @a even odds, the value it gives any claim nothing speaks to, and in the half of the worlds where @a holds the line brings @c to 0.9 (the floor of 4.1). A premise on the negated side, $e 0.9 @c | ~@a, stays at 0.500 as well (both measured 2026-09-24 with solve_map.py at its default; the first is examples/toys/a10-t2.argmap, group ::free, in the app the Semantics audit map audit-reference, group R1).

That 0.500 is a placeholder. Nothing in the map says what the premise is worth, so the solve leaves it at the network’s fill and lets the lines that use it settle it, which is rarely what an author means by a premise. The checkers name every such statement with the info note I7, an unpriced ground (10.1), and the fixes are ordinary authoring:

  1. Author a value on the premise. On a frontier root this is the whole fix: a point or a pair at the source’s register (4.7).
  2. If the source asserts the premise as a claim of its own and you want it discussable as one, give it an attributed premise-less line at that register (the direct-assertion pattern, 7.13).
  3. If the premise was only ever part of the step, fold it into the line that uses it and re-elicit that line’s strength.

A converse is a different thing, and still worth writing. If the source also asserts “and otherwise not” (“nobody builds it, it kills nobody”), write it as a second line on the other side of the premise, ~@c | ~@a. It speaks to the worlds the first line leaves open, the ones where the premise fails, so it moves the conclusion and leaves the premise alone: with a 0.8 converse beside the 0.8 line, @c reads 0.500 and @a stays at 0.500 (the Concepts map c1-drift-tax, block 3). It also keeps the claim on the map, where a reader can argue with it. The flagship map’s $no-doom-otherwise is the worked example (7.7).

What does move a premise is information about what follows from it:

  1. A confirmed consequence raises it: pin @c at 0.9 beside the lone line and @a reads 0.593, and each further confirmed consequence raises it again (the fork gallery’s f-consilience, MATH §5).
  2. A refuted consequence lowers it (modus tollens; f-refutation).
  3. A support and an objection on the same premise whose strengths sum past one make their shared case rarer: 0.8 against 0.3 reads @a at 0.466, where 0.8 against 0.15, which fit, leave it at 0.500 (4.4).

Each is the map’s own inference about the premise, and the residual rule (4.3) says to leave it standing.

Until 2026-09-08 this section taught the drift tax: under the retired uniform reference a lone 0.8 line dragged its free antecedent to 1/(1+2^0.8) = 0.365, a line on the negated premise pushed it up to 0.651, a confirmed consequence lowered it (0.464 where the shipped solve reads 0.593), and the remedies were ways to cancel that pull. The network reference D161 ships has no such pull; MATH §4.4 keeps the record, and solve_map.py --reference d36 reproduces every number of it (the same scratch toys, 2026-09-24).

4.3 Statement values and the residual authoring rule

A statement’s authored value is a floor-style constraint, and the single most important discipline in the whole system applies to it:

Author only the evidence for or against a statement that is not already contained in the rest of the map.

  1. Frontier roots (statements with no incoming evidence in the map) keep their authored values. Their number is the map’s interface to everything unmapped; that is what roots are for.
  2. Derived statements (concluded into by mapped evidence) should normally carry no authored value. Their probability is the output of the solve. If you author one anyway, you are counting the mapped support twice.
  3. If you disagree with what the solve delivers for a derived statement, you have three honest moves, in order: fix the argument (structure or strengths); add the missing evidence as a new, named line (a premise-less evidence is fine, but it must say what the evidence is); or record your number as a check credence, # check: p, and let the displayed badge show the disagreement.

The check credence is the designated home for “all things considered I believe 0.85 even though the mapped argument delivers 0.48”. It is displayed, compared, and never constrains the solve. Wanting to force a derived statement to a number is precisely the situation the rule exists to catch.

One explicit anti-pattern: a statement line carrying both an authored value and a # check: comment. On a concluded-into statement the pin double-counts the mapped support and, worse, fights the very evidence you authored against it (the pin holds the solved value where the counter-evidence should have moved it), while the check silently disagrees with the pin. Concluded-into statements take a check or nothing; only frontier roots take pins. The lint and the editor both flag a bare point beside a check as W23 (since 2026-08-08, narrowed 2026-08-22 to the bare point), on any statement line, because it is wrong wherever it appears: one extraction wrote both on 28 nodes, and the only visible symptom was that the gaps looked suspiciously small. A bare point beside a strengthed line that concludes into the statement is W25 (2026-08-22), whether or not a check sits beside it. An explicit residual pair beside a check (0.6?/0? with # check: 0.9) is silent: that is the intended two-slot shape, the pair being the unargued remainder and the check the total (4.7).

The rule is topological and does not change inside refinement boxes. A hinge statement that sibling lines inside a refinement conclude into is a concluded-into statement like any other: check, not pin, even when the source asserts it at a clear register and the mapped internal support delivers less. That under-delivery is an audit finding about the source, not a display problem to pin away. When the source asserts the hinge directly, over and above the arguments it gives for it, that assertion is itself evidence and has a named home: a premise-less, attributed evidence line at the source’s register (the direct-assertion pattern, 7.13). It accumulates with the argued routes instead of clamping over them, and it is visible and criticizable in a way a pin never is. Under v0.3, a floor pair is the interval-shaped variant.

The distinction that keeps this straight: a statement’s own indented block explicates its number (the block refines the implicit factor asserting the marginal, and the statement keeps it, 3.4); sibling lines concluding into the statement replace it (the value becomes the solve’s output, and the author’s number moves to a check).

Also: inference through the map is not evidence you authored. If an evidence about @a -> @c moves the solved P(@a) (a confirmed or a refuted consequence, 4.2), do not “correct” @a’s authored value for it; that effect is already contained in the map.

A root’s number has a second axis too (D161, 2026-09-08): how firmly the solve holds it under a reader’s what-if comes from the width the author left open, an interval 0.7/0.1 about 8 flips and a point about 200, the cap. Nothing is elicited for it; the width you write for the register (4.7.4) is the firmness. Section 4.7.2 gives the rule.

4.4 Independence, and what to do when it fails

Separate evidence lines are treated as independent mechanisms; their premise-less masses accumulate like independent reasons (noisy-OR). That is what makes two convergent lines mean “two independent reasons”. When the grounds actually overlap, independence double-counts. Three repairs, in increasing order of structure, follow the next paragraph.

Independence holds among the lines on one side. A line for a statement and a line against it that apply to the same case are read together (since 2026-09-23, MATH §3.8): they are shares of one population of cases, so they never fire in the same world, each keeps the share its author gave it, and independence is what gives where the numbers do not fit. A granted objection therefore caps what the supports can claim in the cases where it applies, however many supports there are: five 0.9 supports beside a 0.15 objection read 0.90 in the case where all their grounds hold, where weighing them as independent evidence read 1.00 (measured 2026-09-23 on a seven-statement toy with every ground free). To move an objection, answer it with an undercut (4.5), lower it, or doubt its grounds. Where the strongest support and the strongest objection sum past one, the overlap is a contradiction that makes the case rarer, and the pressure lands on the grounds that produce it. If the two lines really describe different cases, name the statement that separates them, and they stop meeting.

The three repairs for overlapping grounds:

  1. Merge the lines into one evidence if they are really one argument.
  2. Name the shared source as a statement and condition both lines on it; the dependence is then authored in the world layer where it belongs.
  3. Complementary partition: make an “even if” explicit by conjoining the negation of the other route, as in $mwb-time ... | @wont-solve-in-time AND ~@align-hard (the two routes then partition the worlds instead of overlapping).

One consequence of independence is worth meeting before it surprises you. Two independent reasons for one conclusion stop being independent once the conclusion is known: with the effect settled and one cause observed, the other cause is less likely, because the effect is already accounted for. That is explaining away, and the solve shows it, as any probabilistic network does. D161 declares it as a divergence from the conditional-logic literature’s syntax-splitting postulate, which would hold the other cause where it was; the map takes the Bayesian side. The exhibit is examples/toys/f-explaining-away.argmap (in the app, the Semantics audit map audit-forks, group F6): two causes each bring the effect at 0.9; with only the effect observed both causes read 0.573, and observing one of them as well drops the other to 0.515 (measured 2026-09-08 with solve_map.py at its default on the toy, the same numbers the kernel is pinned to in the solver-compile toys test; the record’s hard-mode figures are 0.573 and 0.512). Nothing needs authoring around it. A reader who pins one cause of a known effect in what-if mode and watches the other cause fall is seeing this.

Two boundary clarifications. First, the discipline applies to lines converging on the same conclusion; one statement legitimately feeds premises of several different conclusions, and that needs no declaration. Second, genuinely independent routes stacking a hub high (noisy-OR takes four 0.85 routes past 0.99) is not by itself an error: if the source really asserts four independent sufficient reasons, the high number is what the source’s own logic delivers, and a lower check credence on the hub turns the difference into a visible audit finding (the source claims less than its own arguments compound to). Before accepting that reading, check whether the routes share an unnamed latent (repair 2); several “distinct” failure modes of one mechanism usually do.

4.5 Undercut strength

An undercut (3.6) carries the defeater’s operative rate: granted the grounds, how often does the targeted inference actually fail? The grounds’ own plausibility is carried by the guard statements, so do not pre-discount the undercut for it. Likewise, do not pre-discount a defeater because a response to it exists; author the response as its own undercut of the undercut and let the graph do the discounting. q’ = 0 is inert, q’ = 1 eliminates the target in context, values between interpolate.

An undercut does not, however, push its own conclusion. This paragraph said the opposite until 2026-07-27, when the skill-only sufficiency eval authored a cluster on the strength of it and produced a 0.61 statement gap; the claim is wrong and the correction matters for authoring. An undercut-shaped line compiles as a pure inhibitor of its target: per the factored-A compile (SOLVER_SEMANTICS §1.2), “inhibitors carry no zero-set of their own”, so the negated conclusion the line names receives no independent floor from it. Measured under the shipped solve (2026-09-24, solve_map.py at its default): an undercut whose target is unstrengthed leaves its conclusion at 0.500, exactly as if the line were absent. And an answer to an objection, an undercut of a rebuttal, recovers the claim toward the value it would hold with the objection absent and never past it. The flagship’s shape, in four statements:

---
title: an objection and its answer
argmap-format-version: 0.3
---
@s [S] 0.8: the support's ground
@w [W]: the objection's ground, left open
@g [G] 0.85: the answer's ground
@c [C]: the claim
$sup [S supports C] 0.9 @c | @s:
$obj [W argues against C] 0.2 ~@c | @w:
$ans [G answers the objection] 0.9 @c | @g AND $obj:

C reads 0.859 with the support alone, 0.823 once the objection stands unanswered, and 0.840, 0.852 and 0.855 with the answer at 0.5, 0.9 and 1; even at 1 the answer leaves the objection standing in the 15 percent of cases where its own ground fails. Take the support away and the objection with its answer reads 0.488: above the objection alone (0.450), below the 0.500 of a claim nothing speaks to. The answer revives the worlds the objection killed and delivers nothing of its own. The file as shown is the toy examples/toys/u-grounds.argmap, group ::guard-sup with its ids unsuffixed, and the supportless one its group ::guard (in the app, the Semantics audit map audit-undercuts, group U3); the other readings are the same lines with one changed or removed. T13’s reinstatement sweep (u-t13, audit-undercuts U2) reads the same way on the inference itself. The figures this paragraph quoted until 2026-09-24 (0.500 to 0.866 against a rebuttal-free 0.898) were measured in July under the retired uniform reference.

The authoring consequence: when the source both raises an objection to an inference and asserts the fact that objection rests on, the undercut carries only the first. If you want the fact to bear on the claim as well, give it its own ordinary evidence line beside the undercut. That is not double-counting (the inhibitor acts on the inference, the plain line acts on the claim), and without it the fact the source actually reports is silently absent from the solve. The same holds for an answer’s ground: wired as its own 0.9 line into C beside a guardless answer, the ground brings C to 0.879 where the same ground inside the answer’s guard gave 0.488 (u-grounds, ::split against ::guard, measured the same day).

Which objections are undercuts at all (D170, 2026-09-28). An undercut says a step is unreliable; where it holds, it says nothing about which way the claim goes, so a reader who wins an undercut outright is left with cases no link speaks to, and those count as even odds. Before wiring an objection, ask what is true instead if it is right. If the answer is that the step cannot be trusted, it is an undercut. If the answer names a claim that is false, it is an ordinary line into that claim’s negation, and the answers that say its inference fails stay undercuts of it. On the flagship, “experts disagree” stays an undercut of $link-fine, while “if building killed the builders, they would stop” and “nothing so far has ended us” are lines against @mis-ext. Won outright, each now takes the title claim to 0.10 through $harmless-no-extinction; as undercuts, the first left it at 0.37 (measured 2026-09-28 with solve_map.py at its default).

Readers arriving from probabilistic conditional logic: the translation rules under 3.6 say when a low conditional is an opposed line and when a subclass exception needs this undercut plus a rebuttal.

4.6 Solved values, tension, and the what-if

The display is computed-first: the solved value is the primary number, the authored values are the claims it was asked to honour, and tension is the distance from the solved value to what the author wrote. Since D161 that distance is measured against the whole of what was written: a point, an interval’s two bounds, a step’s strength as a lower bound. A tint means “the map could not honour this number, by this much”; the solve settling somewhere inside an interval the author left open is a readout without colour (2.2 has the rule, the 0.01 tolerance and the flagship’s counts). A tension badge is information about the argument, and it stays: the flagship’s widest, the headline 0.084 short of its check, is an audit finding about the source (2.2).

Authored 0 and 1. Since D161 a point is held at the point cap, about 200 flips, so an authored 0 or 1 no longer deletes possible worlds (it did under D36: “world-killers”, SOLVER_SEMANTICS P3). It still claims a certainty the source rarely states, and it is held no more firmly than 0.97 is. If you mean “very confident”, write 0.97, or give the interval as a v0.3 pair. The honest wide statement is cheap; the false point is not.

The what-if is revision. In what-if mode a reader moves a claim’s number. What the map does with it was ruled on 2026-09-08 (D161, after the two-verb discussion in ideas/plans/semantics-antecedent-pull-2026-09-06.md section 15), and it is one rule with two parts:

  1. The reader’s number replaces the author’s on that claim (the author’s point or interval there is set aside for the duration) and enters as a point at the cap: your number is a claim like the author’s, as firm as a point (198 flips, the cap). A reader’s 0 or 1 is therefore a very firm claim, never a fact.
  2. Everything else the author wrote stays in force, each number at its own firmness, and the map re-solves. The question answered is: if this number were as the reader says, which of the author’s other numbers gives, and by how much.

Two cases follow, and both are worth knowing. The numbers are the flagship’s under the ruled counts (width counts on its intervals, kind counts on its lines with deductive at 1000, the reader’s row at the cap; re-measured 2026-09-23 on the map as it now stands with solve_map.py under the shipped solve, the one draw of MATH §3.8 included; the 2026-09-08 figures this section quoted before came from an earlier state of the map and the rule before the one draw):

On a root, the map follows you forward. @llm-nice (the authors’ 0.85/0.05 that current systems seem nice) moved to 0.1: the map’s answer is @nat-abs 0.20 to 0.11 and nothing else past 0.01; no line bends further and nothing new colours. The reader changed a premise the author argued nothing for, so there is nothing to give; the map propagates. A better-connected root shows the same at scale: @instrumental-convergence to 0.2 moves eleven statements past 0.01 (@incorrigible 0.83 to 0.57, @not-preserved 0.89 to 0.77), lights one of the book’s own check badges downstream (@not-preserved at -0.18: a conclusion the reader’s premise no longer delivers) and bends three hope lines, the hopes that argue against that conclusion ($c5-selfish-obj 0.70 to 0.63, $c5-digital-obj to 0.67, $c5-leave-obj to 0.68). Every tint sits downstream of the edit.

On a conclusion, the author’s case retreats where it is softest. @everyone-dies (the headline; no authored number, check 0.88..0.98 when these figures were measured, re-read to 0.75..0.95 on 2026-09-25) moved to 0: the map meets the reader (0.0005), bends no line past 0.05, and gives on the free premises upstream, @if-built 0.88 to 0.53 (its check badge lights at -0.32), @asi-soon 0.89 to 0.70, @mis-ext 0.97 to 0.88. Read it as “then it was not built, or not soon”: the book’s case is softest at its premises, because they are derived claims nobody pinned. Pin those too (the record’s P6 set: built, misaligned when built, the misalignment extreme, the link fine, and still nobody dies) and the retreat has to land on the lines and on the authors’ flat assertions: the analogy response $resp-exp 0.90 to 0.63, the testimony undercut $uc-experts 0.30 to 0.54, the mechanism objection $c12dr-obj 0.15 to 0.21, a second analogy $uc-precedented 0.30 to 0.35, and two 18-flip assertions leave their intervals (@precedented 0.11 to 0.25 against 0.05..0.15, @immature-field 0.89 to 0.83 against 0.85..0.95); no deductive line moves, and no hope line does either, because none sits on the path: the count orders what gives among the numbers the reader’s claim reaches, which is why a 64-flip mechanism bends here while the 4-flip hopes stand. That list is what “the book’s case holds least firmly” means, and the Most moved panel shows it with each line’s kind. Since 2026-09-28 the two objections that deny @mis-ext argue against it instead of doubting the step (D170, section 4.5), so the step keeps one doubt to bend, and the same set’s retreat lands on $uc-experts, 0.30 to 0.98, and $resp-exp, 0.90 to 0.27, alone.

The refused residual. The reader’s number is met to within 0.002 in every case above (through the shipped compile and kernel the P6 set’s points read 0.0020, 0.9981 and 0.9982 for @everyone-dies, @if-built and @mis-ext, D161 item 4; @llm-nice at 0.1 reads 0.101, both re-measured 2026-09-08; since 2026-09-28 the P6 points read 0.0044, 0.9956 and 0.9956, one undercut being left to bend), and that is the usual outcome, because 198 flips outrank everything on the flagship except its 69 deductive lines. Where the map cannot meet it, the adjustment row says by how much. On the confusions map’s Socrates syllogism (c4: two premises at 0.99 and a formal step at 0.99) a reader’s “not mortal” at 0 is held at 0.32, the step holding at 0.99 and each premise giving 0.32; re-keyed deductive, the step gives 0.06 and the reader’s 0 is held at 0.30 (4.1, rule 1; both re-measured 2026-09-08 through the shipped solve). The residual is the honest answer, and it is tinted only past the 0.01 tolerance.

The other reading of a what-if, “suppose it turned out that way”, is the conditioning verb: the reader’s number as a fact about the world, the author’s numbers updated by Bayes. It stays a bench readout (solve_map.py --condition, 2.3). Under conditioning the author’s pins would yield; under revision they hold at their firmness and the reader sees which inference they thereby reject. The reason for the choice (recorded in semantics-antecedent-pull-2026-09-06.md section 15 item 12): showing a reader where an argument’s inconsistencies lie is the point of the tool, and conditioning hides that by moving to the unlikely world without showing how surprising it was. One more effect to expect in what-if mode: pinning one cause of a known effect lowers the other causes (explaining away, 4.4).

4.7 The epistemic delicacies: residual, point, pair, check, derive

The rules above keep a map lint-clean. The rules in this section keep a lint-clean map from counting one consideration twice, which is the error class no lint can see in full and the one an extractor has to get right in a single pass. They were settled in August 2026 (DECISIONS.md D152, the register to pair table in AUTHORING_NOTES 2026-08-23, the check intervals of the same day) and each of them has a five-line measurement behind it. Those measurements are the toy battery in examples/toys/; a reader-facing selection of them ships in the picker’s Semantics audit group (examples/audit/, generated from the toys with the measured answers in the group heads). Read the toy, then the rule; the rule is what the number says.

Two measurement dates sit in this section, and they are marked. The T1 to T4 illustration in 4.7.2 was re-measured on 2026-09-08 under the network reference (D161, the shipped default since that day) with solve_map.py at its defaults, as authored and under a reader’s what-if; the toys’ README carries the same numbers in a dated block. The pictures from 4.7.5 on were re-measured on 2026-09-08 the same way (solve_map.py at its defaults; the README’s second table), and each names the retired D36 figure it replaced beside the network reference’s. The rules stand under both references, because they are about what a number is evidence for, which no reference changes; what the reference changes is where the double count becomes visible, and 4.7.2 says where. One picture did not survive the change and says so: the ridge price of a p = 1 terminology statement (4.7.6, T7) was the uniform reference’s, and under the network reference the statement holds at 1.000 and the junction pays nothing.

4.7.1 The residual rule, stated once more

A statement’s own number is evidence not already in the map. Two cases follow:

  1. A frontier root (a statement no strengthed line concludes into) keeps its number whole. Nothing in the map argues for it, so its number is the entire interface to the unmapped world, and the all-things-considered reading is the right one.
  2. An interior statement (at least one strengthed line concludes into it) may carry a number only for the part of the author’s credence that the mapped lines do not trace: the unargued remainder, which the reader-facing vocabulary calls the author’s tacit grounds. The author’s total belongs in the check (4.7.2).

The asymmetry is the whole rule. A root’s number and an interior statement’s number are different kinds of object, and the mistake that produced the 2026-08-03 incident (AUTHORING_NOTES) was writing the same all-things-considered figure into both slots.

4.7.2 Point, pair, check: three slots on one statement

A statement line has three places a number can sit, and they mean three different things.

  1. The point, @s 0.9?. Since D36 a bare point is read as the degenerate pair 0.9?/0.1?, a zero-width interval: the members sum to 1 and P(s) is held at 0.9. Since D161 it is held at the point cap: a point counts as an interval with silent mass 0.01, 198 flips (about 200), so no two points can make the solve infeasible, and a point may solve up to about 0.005 off its value, inside the tint’s tolerance. The cap sets the firmness only; the target stays at 0.9 and no interval is written for it. (Under D36 the point was a spring at weight 400 that leaked about 1/400.)
  2. The pair, @s 0.6?/0? (v0.3, D53). The author’s direct evidence about the statement: a floor of 0.6 for it and nothing against it, so P(s) is bounded to [0.6, 1]. The map’s own inference then selects within the interval: every line the statement feeds or is fed by pushes the solved value around inside it, and the reference picks the point when nothing pushes (under D36 the balancing prior’s maximum-entropy point; under D161 the map’s own network, where the lines put it, and on the flagship every one of the 224 intervals is met inside its bounds, 111 of them more than 0.05 above the floor, 2.2). Inference through the map never counts against the pair; the pair is read as direct evidence and the width is the author’s honest spread. Since D161 the width is also the firmness: by Walley’s imprecise Dirichlet model with prior strength 2, an interval with silent mass m is held at 2 (1 - m) / m flips, so 0.6?/0? (silent 0.4) is 3 flips, 0.7/0.1 is 8, 0.85?/0.05? is 18 and the cap’s 0.01 is 198. A wide interval is an honest spread and an easy give under a reader’s what-if; that is one fact, said twice.
  3. The check, # check: 0.85..0.95. The author’s total, all things considered, as an interval at the register’s width (4.7.4), or a bare point as the zero-width case. It never constrains the solve, so it has no firmness either: under a reader’s what-if a check-only statement is the first thing to move, and its badge says by how much (4.6’s @if-built). The badge a reader sees is the signed distance from the solved value to the interval, zero inside it. A badge that stays is a finding about the argument; closing one by moving a number is the move the whole discipline forbids.

The three toys T1, T2 and T3 put the three slots on the same five-line map. @want is a frontier root, @subvert is the statement whose authored 0.9 silently contains “because it wants things”, and @resists reads @subvert.

---
title: T1 pin kept beside a mapped argument
argmap-format-version: 0.3
---
@want [The system wants things] 0.9?: a frontier root
@subvert [It routes around oversight] 0.9?: the author's all-things-considered 0.9, kept as a PIN
$sub-ev [wanting means routing around obstacles] 0.9? @subvert | @want:
@resists [It resists being corrected] : # check: 0.7
$res-ev [subversion implies resisting correction] 0.85? @resists | @subvert:

The numbers below are the network reference’s, measured 2026-09-08 with solve_map.py at the shipped defaults, first as authored and then under one reader’s what-if (--override want=0.1: the reader’s number replacing the author’s on the root, 4.6).

T1 is the shape to avoid, and the lint says so (W25 on @subvert): the point sits beside a strengthed line that argues for the same statement. As authored it solves @want 0.899, @subvert 0.900, @resists 0.883. The wanting consideration is counted twice here, once inside the 0.9 and once through $sub-ev, and the as-authored solve shows no sign of it, because a pin mostly restates itself.

T2 derives the statement: the same file with @subvert’s head line replaced by @subvert [It routes around oversight] : # check: 0.9. As authored it solves @want 0.899, @subvert 0.905, @resists 0.884, the same picture as T1 to within 0.005. Under the network reference the one line the map wires delivers the author’s total by itself: a 0.9 line under a 0.9 premise fills its conclusion at 0.9 × 0.95 + 0.1 × 0.5 = 0.905 (the F1 fill), so the check is met and no badge lights. That is the honest state too. The badge is the distance between what the mapped lines deliver and what the author holds, and here they agree; it lights the moment they stop agreeing, as the what-if shows next. (Under the retired counting reference T2 read @subvert 0.868 against the check and @resists 0.866 against T1’s 0.879, a badge and a drop that earlier versions of this section presented as the double count made visible; both were that reference’s antecedent pull and went with it, D161 item 1.)

Where the double count shows under the network reference is the what-if. With the root @want set to 0.1, T2 follows its premise down: @subvert 0.545 against its check of 0.9, the badge lit at -0.36, and @resists 0.732. T1 does not move: @subvert 0.899 and @resists 0.882, the pin holding at its 198 flips against a 16-flip line, exactly as it held against the map. A pinned interior statement is deaf to its own premise, and that is what counting a consideration twice means: a reader who doubts the wanting is told it makes no difference to the routing the author derived from the wanting.

T3 adds a residual floor pair beside the check: @subvert [It routes around oversight] 0.6?/0?: # check: 0.9. As authored it solves @subvert 0.896, @resists 0.881, and the lint is silent, because a non-coinciding pair beside a check is exactly the two-slot shape D36 item 3 describes. Under the same what-if it reads @subvert 0.840 and @resists 0.857: the floor absorbs most of the reader’s move (the map now explains @subvert through the floor instead of through @want), which is fine when the 0.6 is the unargued remainder the text licenses and is the pin again with extra steps when it was read off the total. T3’s hazard is elicitation; the machinery is sound.

T4 is the legitimate interim state: the pin kept, the relation drawn, the strength left off ($sub-ev [..] @subvert | @want:). An unstrengthed line compiles inert, so T4 solves as the pin alone (@want 0.899, @subvert 0.899, @resists 0.882), the lint lists the line under I3, and nothing is asserted twice. Against T1 the strengthed edge beside the pin is worth 0.001 as authored and nothing under the what-if (T1 and T4 both answer @subvert 0.899, @resists 0.882 at @want 0.1): next to a point at the cap, a 16-flip line adds nothing the pin had not already asserted. (Under the counting reference the lift read +0.006 and +0.002 and stood here as the double count visible even in a toy; under the network reference the count makes the pin outrank the line outright, and the visible sign is the refused what-if above.)

4.7.3 Deriving a root (D152)

To derive a root is to give a pinned root its first incoming strengthed line. One test decides whether the line may be written, and one obligation follows once it is.

The inference test. The line is written when it carries an inference (premise to conclusion) and withheld when it would only restate, or co-refer to, the same episode or proposition at a second dock. A restatement wired as inference counts one consideration twice. The flagship’s precedent is @psychosis and @c13ws-retrain: one episode family at two docks, cross-referenced in comments and never joined by a line. Source faithfulness is no part of the test: a real inference the source omits is a comprehensiveness defect of the source, and it is mapped, with a comment marking it as the mapmaker’s.

What does not count against a derivation. “A heavily consumed root, once derived, moves everything downstream” is true and is no reason to withhold the line: the movement is the coherence audit working. Completion is always licensed; if a correct completion swings the conclusions wildly, that refutes the semantics or the elicitation and never the completion. Measured on the case that settled it: deriving @steering-finds-subversion from the two grounds the book itself gives (| @goals AND @deep-gears) moved that statement 0.893 to 0.628, @goals 0.842 to 0.712 and @incorrigible 0.821 to 0.695 (the two-floor pass, AUTHORING_NOTES 2026-08-23, under the retired uniform reference, whose antecedent pull amplified every such move). Under the shipped solve the same derivation moves it 0.899 to 0.843 and @incorrigible 0.856 to 0.834, and leaves @goals at 0.851 (the map as it stands against a copy with the pin restored and the line removed, solve_map.py at its default, 2026-09-24). Either way it is a finding about the book (its flat 0.9 outruns its own two grounds), and the map now shows it.

The elicitation obligation. Once a line sits under the root, its authored point is no longer a residual. The point moves to the check, at its register’s interval, and the residual slot is empty by default for every register: a floor read off the register is read off the total, because the register describes the all-things-considered credence and the text never gives the split. Two text-licensed exceptions, applied row by row:

  1. The passage names a second ground the map does not yet wire: a floor at that ground’s own register, with a gloss naming the passage, as the interim encoding until the line is written.
  2. The passage says the listed grounds are a subset (“to name a few”): a remainder at the fallback row 0.2?/0?, with the quote in the gloss. Small enough never to close a badge; nonzero so the acknowledged mass shows.

In practice the derivation is one head-line edit plus one new line:

# fragment - not standalone (flagship, anchor "$subversion-ev"; before and after the 2026-08-23 derivation)
@steering-finds-subversion [A goal-directed system seeks ways to subvert whatever limits it] 0.9?: shared ground (reused by Ch5, Ch11)  # C7 flat: the state before
@steering-finds-subversion [A goal-directed system seeks ways to subvert whatever limits it]: shared ground (reused by Ch5, Ch11)  # check: 0.85..0.95 (C7 flat; derived via $subversion-ev)
$subversion-ev [wanting routes around obstacles the gears can find] 0.9? @steering-finds-subversion | @goals AND @deep-gears: the two grounds the authors give

(The fragment shows both head lines for the comparison; a real file carries one of them.) Before writing the edge, grep the map for the root’s id: the reason it was left unwired is usually written down in a comment somewhere, and a derived root’s comment has to be rewritten along with its number.

4.7.4 The register to pair table, in brief

The source states neither points nor intervals; it states registers (flat repetition, “probably”, “our best guess”, “nobody knows”). The D39 rubric mapped registers to points; the 2026-08-23 table adds the pair column, pre-registered before any root was touched. Its rules:

  1. Centring. A pair’s midpoint (1 + p⁺ − p⁻)/2 equals the rubric point, and the register fixes the width: p⁺ = point − w/2, p⁻ = 1 − point − w/2. Widths are rubric constants: 0.05 for the strongest categorical class, 0.10 for flat assertion, 0.20 for “by default” and “best guess”, 0.30 for “we expect” and “could well”, 0.40 for the weakest hedge, 0.80 for a refusal. Authored mass is linear probability; log-odds widening was only ever a sweep’s parametrisation.
  2. The rows, point to pair, for a root that stays a root:

    register (flagship class) point pair
    repeated categorical, anti-hyperbole (C1) 0.97 0.95?/0?
    categorical deontic, “insane gamble” (N1, N3, RS1) 0.93 0.9?/0?
    title register, C2 0.93 0.88?/0.02?
    flat unhedged assertion (C7, P-FACT) 0.90 0.85?/0.05?
    “by default”, “our best guess” (C3) 0.85 0.75?/0.05?
    RS3 0.80 0.7?/0.1?
    “we expect”, a granted point (C4, P-GRANT) 0.75 0.6?/0.1?
    “could well” (C5) 0.70 0.55?/0.15?
    weakest hedge (RS6) 0.60 0.4?/0.2?
    role defaults 0.8 / 0.7 / 0.6   0.65?/0.05?, 0.55?/0.15?, 0.45?/0.25?
    explicit refusal, “nobody knows” (COIN) 0.50 0.1?/0.1? plus a gloss sentence
    the objection the source denies at class X (P-CONTEST) 1 − X the mirror of X’s pair
  3. One-sided categoricals. Where the source leaves no room against (0.9?/0?, 0.95?/0?), the midpoint sits above the point, and that is the reference’s fill (the balancing prior’s under D36, the network’s under D161), never a number the table asserts.
  4. The P-CONTEST mirror. An objection the source denies at class X takes X’s pair reflected (0.05?/0.85? for a flat denial, 0.1?/0.6? for a hedged one), never 0/p: omitting the floor would leave nothing for the objection, and “we don’t think so” is weaker than “impossible”.
  5. Refusals are wide pairs with a gloss. “Nobody knows” where the statement is the unknown becomes 0.1?/0.1? and a gloss sentence naming the passage, in one shared shape: what the passage does (“the page calls the question open”), then “the wide pair carries that, and its midpoint is nobody’s belief”. The interval is the message.
  6. Per-node asymmetric pairs are written where the passage states both directions (an inner flat claim under an outer hedge, a concession, “routes against their own”): from the text, with the rationale in a trailing comment, and the midpoint may differ from the class point. The class row is the default where the text gives one direction.
  7. Riding rules. ? on every member, zero members included; a gloss sentence whenever the total width exceeds 0.30; the class token in the trailing comment of every pin; nothing on an evidence line moves.
  8. The width is the count (D161, 2026-09-08). The same widths fix how firmly the solve holds the pair under a reader’s what-if, by Walley’s rule n = 2 (1 - m) / m on the silent mass m: width 0.05 is 38 flips, 0.10 is 18, 0.20 is 8, 0.30 is 4.7, 0.40 is 3, the refusal’s 0.80 is 0.5, and a point’s capped 0.01 is 198. On the flagship the median pin is 18 flips and the firmest 38, below a mechanism line’s 64: under revision the authors’ assertions give before their mechanisms, which is the completion principle’s order and the story 4.6 tells.

One proviso, stated so it stops being rediscovered: a root’s register-read confidence is a posterior that may already discount what the author believes downstream, and reading it as direct evidence lets the solve apply modus tollens a second time. There is no operational alternative (the text never gives the un-discounted figure), the magnitude at these widths is small and measured (at most 0.043 per statement in the widening sweep, mean 0.006 on the centred table), and the compensation is to show the interval and the band rather than present a midpoint as the author’s number.

4.7.5 Shared considerations

Separate lines are combined as if independent (4.4). The only way the map can say two lines co-vary is a shared statement they both cite, so whenever two lines rest on one consideration, name it as a statement and condition both on it. The sign of the effect tells you whether you had it right. T5 and T5b put one consideration into an AND junction twice, first through a pin and then through the shared root:

---
title: T5b the same junction with both routes derived
argmap-format-version: 0.3
---
@want [The system wants things] 0.9?: a frontier root
@subvert [It routes around oversight] : derived from the same root the other route uses  # check: 0.9
$sub-ev [wanting means routing around obstacles] 0.9? @subvert | @want:
@grabs [It grabs resources] : # check: 0.9
$grab-ev [wanting implies grabbing] 0.9? @grabs | @want:
@danger [It is dangerous] : # check: 0.9
$danger-ev [subversion and grabbing together] 0.9? @danger | @subvert AND @grabs: the wanting consideration enters twice here

With @subvert pinned at 0.9 instead (T5) the conjunction solves @danger 0.866; derived as above it solves 0.876 (network reference, 2026-09-08; 0.843 and 0.850 under D36). The conjunction rises once its two premises are correlated through the shared root, because independence understates a conjunction of positively correlated premises, and the pin had asserted independence.

T8 and T8b are the same fingerprint with both consumers present: two study findings that share one methodology statement @method (T8) or rest on separate @method-a / @method-b roots (T8b), each read by an AND consumer and an OR consumer.

---
title: T8 two lines sharing a premise co-vary through it
argmap-format-version: 0.3
---
@method [The shared methodology is sound] 0.7?: the common cause, named as a statement
@a [Study A's finding holds] : # check: 0.65
$a-ev [A's report, given the method] 0.9? @a | @method:
@b [Study B's finding holds] : # check: 0.65
$b-ev [B's report, given the method] 0.9? @b | @method:
@both [The effect is real, both needed] : # check: 0.6
$both-ev [two findings together] 0.9? @both | @a AND @b:
@either [The effect is real, either suffices] : # check: 0.8
$either-ev [either finding suffices] 0.9? @either | @a OR @b:

The two findings do not move between the maps (@a and @b 0.815 in both; 0.750 against 0.751 under D36). The consumers move in opposite directions: the AND rises (@both 0.799 separate to 0.818 shared; 0.751 to 0.782 under D36) and the OR falls (@either 0.935 to 0.915; 0.918 to 0.886 under D36), because findings that share a methodology stand or fall together, which helps a conjunction and costs a disjunction. argmap-query shared-cause lists every such overlap in a map, one row per shared statement; each row is either the deliberate world-layer device or an overlap to factor out, and the remedy is the same either way: name it.

One shape the static audit cannot see: a statement conjoined with its own derivative. The flagship had $wst-race conclude from @race-dynamics AND @one-cavalier-suffices while $ocs-ev derived the second premise from the first, which counts the race consideration twice through one junction (the covariance probe’s single hit above its floor, +0.029 of manufactured co-movement). The fix is to drop the duplicate premise, or to re-elicit the line conditionally on the derivative alone. Trace each premise’s ancestry before writing an AND.

4.7.6 Definitional versus substantive bundling

A node that bundles a definitional part with a substantive part is two variables in one slot. “A goal-directed system seeks ways to subvert whatever limits it” contains “routing around obstacles is what wanting means” (a definition) and “oversight is one of the obstacles” (the claim). Two remedies: split the node into one proposition per statement, or derive it from both halves, one line per ground (the $subversion-ev shape above). No lint sees this; the authoring rule is the tool.

A definition that does inferential work encodes as p = 1 evidence lines, and a biconditional as two of them, written as the converse pair: the forward line and its contrapositive, both concluding into the defined term (T6, respelled 2026-08-23):

---
title: T6 a definition encoded as a converse pair
argmap-format-version: 0.3
---
@steers [It steers toward outcomes] 0.9?:
@routes [It routes around obstacles] 0.85?:
@want [It wants things] : wanting, by definition, is steering plus routing  # check: 0.8
$def-fwd [definition, forward] 1 @want | @steers AND @routes:
$def-conv [definition, converse] 1 ~@want | ~@steers OR ~@routes:

It solves @want 0.763, which is P(steers AND routes) exactly (0.899 times 0.849; network reference, 2026-09-08; under D36 it read 0.755, within the ridge’s 1/w of the product), and the lint is silent. The same two lines spelled as forward plus backward ($def-back 1 @steers AND @routes | @want) carry the same logic, but they form a directed cycle, so the lint reports W2, and the backward line concludes into the roots, so it draws W25 on each; and since 2026-09-08 the checkers reject the spelling with E11: a line concluding more than one statement is unimplemented under the shipped solve until block normalization lands, and the message hands you the split, one line per statement at strength 1 (the retired D36 solve read it 0.752, the same as the pair). Nothing in the semantics prefers that spelling; the converse pair is the one to write. One-way is not a definition: a single forward line leaves the term free above the members (the reference fills the interval: the balancing prior under D36, the network under D161), so a consumer of the term reads more than the members deliver (in the toy battery’s t9c, @ab 0.896 against the members’ product 0.792 under the network reference, measured 2026-09-08; 0.828 against 0.749 under D36 as pinned 2026-08-22). Both directions, or none.

A pure terminology node is the shape to avoid: @asi-def [ASI means an AI far beyond every human at every cognitive task] 1: conjoined into a premise (T7) draws an edge into every junction that uses it and says nothing a gloss would not. Under the retired uniform reference it also had a price: the p = 1 statement solved 0.988 under the ridge, so every junction it joined paid about 1/400 (@dies 0.849 with the conjunct against 0.854 without, T7b). Under the network reference (2026-09-08) that price is gone: the statement holds at 1.000 and @dies reads 0.859 with the conjunct and without it. The rule stands on the clutter alone. Terminology belongs in a gloss or a glossary surface; only definitional claims become lines, and a definition is a p = 1 line, never a p = 1 statement.

5. From source text to map

This chapter is the workflow that produced the flagship map, distilled from the re-authoring passes logged in AUTHORING_NOTES.md (2026-07-16 through 2026-07-24). It assumes you are extracting an argument from a source (a book, an essay, a debate); mapping your own argument works the same way with yourself as the source.

5.1 Pass 1: skeleton

Extract structure only. Statements, evidences, refinement nesting, labels, glosses, citations; no numbers (unstrengthed lines are legal and compile-inert).

Work in three sub-passes, in this order, because the two extraction directions fail differently: pure top-down invents structure (hub statements the source never asserted, which then solve near-tautologous), while pure bottom-up extracts each section faithfully and never authors the connective tissue: a buried spine (W21) and shared grounds double-counted across clusters. The 2026-08-03 He-essay extraction hit both in one day.

  1. Spine first, top-down: transcribed, never invented. If the source draws its own overview (a section-2 diagram, an abstract’s roadmap, a title conditional), transcribe it: the top tier, the declared foci (focus:), the region list and prefix scheme, recorded as the manifest comment (chapter 8). Where the source asserts its own structure (“this is not a conjunction”, “either route suffices”), that assertion is quotable content and belongs to this sub-pass, not to your judgment. Genre flip: a debate or interview asserts no overview up front; there, extract bottom-up first and write the spine once the meta-shape emerges (usually in the wrap-up). Never fake a spine the source did not assert: an invented spine is the top-down failure mode wearing a checklist.
  2. Regions bottom-up, in source order. Local links under local conclusions, each step citing its sentence; this is where source fidelity lives. The decisions to make deliberately:
    • Statement granularity: what gets to be a claim. A statement label must be a proposition. If a source item is not a premise-to-conclusion step (a meta-principle, a design artifact, pure framing), keep it as a folded gloss or one node with a comment, not as a conjunct in the inference chain.
    • Linked vs convergent: AND only where the step genuinely needs all conjuncts (test: does the inference fail if this conjunct alone is false?). Independent routes are separate evidence lines. If a comment says “independent paths” and the factor says AND, one of them is wrong.
    • What each objection targets. For every objection ask: which inference does this grant, and which does it deny? An objection to an inference is an undercut conditioning on that $id; an objection to a claim is a rebuttal concluding ~@id. Mis-typing this is the most common structural error in first passes.
    • Objections travel with their answers. A response cites its objection, never the reverse, so a map loses rebuttals more easily than it loses attacks: skim extraction, later trimming, and source prose itself (objections are stated loudly, answers quietly) all bias toward attacks left standing unanswered, and the solve then prices an unanswered attack at its full authored strength. Measured on the settled-question benchmark (the H. pylori map, experiments/solver-prototypes/GROUND_TRUTH_PROBE-2026-08-13.md): deleting map lines at random flipped the known-true conclusion in a majority of orders, and the flips were driven by responses dying before their objections. So when you record an objection, hunt for the source’s answer with the same diligence you gave the objection, and when you must cut for scope, cut the objection and response as a pair rather than the response alone. The answered-attack audit (argmap-query, reference in mvp/README.md) lists every attack and who answers it.
    • Nest each cluster’s internal traffic (grounds, caveats, objection pairs) under its target.
  3. Reconcile: where the real work is. Shared grounds are discovered in sub-pass 2, not planned in 1; promote each to the home the burial test picks, which is the nearest container covering every consumer, not blindly the document top (a ground consumed only inside one case block homes at that block’s top rank). Merge or partition lines that turn out to share grounds (chapter 7’s overlap repairs). Set the tier per chapter 8’s spine test: the source’s disclosure order is the guide (the He map’s depth 0 is its abstract, depth 1 its section-2 overview, depth 2+ its detail sections), with cross-tier edges kept visible as top-level evidences or coarse hulls. Then run the lint, and the nest-audit readout for fold candidates you missed.

5.2 Pass 2: numbers, blind, by rubric

The numbers should reflect the source’s confidence, not your own and not what makes the map solve nicely. The discipline that keeps this honest (pre-registered for the flagship map as D39):

  1. Fix a verbal-to-probability rubric before assigning anything: a table from the source’s confidence language to values. The flagship rubrics (AUTHORING_NOTES 2026-07-19) map, for example, categorical repeated assertions to 0.93, flat unhedged entailments to 0.9, “by default” claims to 0.85, “could well” to 0.7; grants of an opposing point take the conceder’s register. Keep two tables, one for statement registers and one for inference-step language (the flagship’s statement classes vs R-STEP), even if the values happen to coincide. Where the source is silent, a role default applies: a value your rubric assigns to a structural role rather than to any phrase (a default for an unhedged asserted step, one for an objection the source raises to deflect, and so on); define these in the rubric itself, because you will need them. Convention: the rubric lives in a comment block immediately after the frontmatter.
  2. Assign all values before the first solve, and do not move them afterwards. If the solve surprises you, the finding is about the argument (or the rubric), and it should be recorded, not tuned away.
  3. Mark provenance. Every rubric-derived number carries ?. A bare number is reserved for values the source states as a credence or probability (“ten to twenty-five percent extinction odds”). A stated frequency or rate (“below one fatal accident per twenty million flight hours”) is not a credence: keep it in the gloss and derive the statement’s value from the assertion’s register as usual. A stated number on a derived statement goes in a trailing # check: comment, never a pin (rule 4.3.2).
  4. The composition rule when hedges stack: the outermost hedge governs. When two readings are defensible, author the weaker and log both. The same rule covers a source that asserts one proposition in two places at two registers: author the weaker register, note the stronger in the gloss or log.

5.3 Pass 3: review

The seven correction classes that were actually needed, in review- checklist form (every one was discovered as a correction, not foreseen; AUTHORING_NOTES 2026-07-19):

  1. Strength provenance. Does every strength trace to source confidence language through the rubric? Is ? on everything rubric-derived?
  2. Undercut target typing. Per objection: which inference does this grant, and which does it deny? (Policy objections wearing implication-undercut shape were the flagship’s most instructive mis-typing.)
  3. Overlap double-counting. For every same-polarity convergent pair: merge, factor out the shared span, reroute an instance-of to the shared ground, or leave independent and say so in a comment (silence is indistinguishable from an unaudited pair).
  4. Connectivity. Every non-headline statement should feed some evidence. Dangling sub-conclusions are usually missed links. A norm the source argues for should not stand unargued in the map.
  5. Nesting. First passes come out flat. Fold clusters under their local conclusion; keep cross-cluster shared ground at top level.
  6. Coverage. Do a full-source pass before calling the map faithful; record the source’s scope in the frontmatter. Summarizing from memory under-extracts.
  7. Mechanical smoke. Run the lint, the real parser (open the file in the editor), and a headless solve (chapter 10) before calling it done.

Two additions from later passes:

  1. Defeat presupposition. For every response/rebuttal: which epistemic state does its ground presuppose? If a response only works while X is undemonstrated, conjoin the statement that says so (a guard), so the defeat lapses in worlds where X is demonstrated.
  2. Multi-voice overlap. When two speakers concur non-diametrically, do not give them independent convergent lines. Full concurrence is one line at the weaker register; a subset relation is the shared span plus a residual increment elicited conditional on it; an instance supports the shared ground, not the downstream conclusion; genuinely disjoint mechanisms stay independent with a comment saying so. A joint line quotes both speakers saying the sentence: if its gloss has to argue that one of them concurs, he has not, and a grant of the other speaker’s point is joint only when the granted sentence is also that speaker’s own words on the map (a grant of a program the other never states as that sentence is the granter’s line). Read a quote to its full stop before any part of it carries a line; a sentence cut at a comma carried a joint line for one afternoon on the debate map before its second half put it back in one voice (AUTHORING_NOTES 2026-09-17 and 2026-09-18).
  3. A spoken refusal to price. On a multi-voice map, when a speaker says on the record that they will not put a number on a claim, put their words on that statement as a > quote line and list the claim under their key in the frontmatter’s declines: block (section 3.9’s display hints, D164). Their view then shows the refusal where a number would stand. Without the quote on the statement the entry is rejected.

5.4 Quoting and citation discipline

Keep verbatim spans to at most one sentence, roughly 25 words, normally one per node; never alter a quote silently; never reproduce a self-contained creative unit (a parable, a poem) whole; retell and compress instead. Whole-map verbatim budget from any single work: low hundreds of words, and for a short source proportionally less (a tenth of the source is far too much regardless of the absolute count).

Where the span goes is the 3.12 test. A quote that supports the claim belongs on its own > line with a [^ref] locator; a verbatim phrase that is a grammatical part of the gloss sentence stays in the gloss in plain double quotes, optionally echoed by a > line carrying the full sentence. The budget above counts both. (The older convention of marking in-gloss quotes ~"…" is retired, and W20 flags any survivors.)

6. Labels and glosses

Statement labels and evidence labels do different jobs.

  1. A statement label is the claim itself, a proposition, and may be a full sentence. Statement labels are not length-linted.
  2. An evidence label states the step. It names its subject and says, in plain words and one clause, what the premises give the conclusion: “a tiny target and imprecise training make alignment hard”. (Sharpened 2026-09-25, Felix’s ruling on the flagship’s @align-hard box; AUTHORING_NOTES 2026-09-25, the label pass. It read “the headline warrant, in a phrase” until then.) The reason is the layout: a line sits where the eye lands first and its premises in a stack the reader skips, so the label is read before its premises, and cold. Four consequences:
    1. Never the warrant alone as a fragment (“the hardness route”, “concedes the abstraction, denies the rescue”), and never a bare “it” whose noun is on another card (10.4 item 12).
    2. No figure of speech unless it is the source’s own image and the gloss unpacks it (“a tiny target, aimed at with blunt tools” was neither).
    3. Never a restatement of a premise’s label (“what can be built, gets built” under the premise “What physics permits, someone eventually builds”): a reader who skipped the premise gains nothing, and one who read it reads it twice.
    4. Every strengthed line gets one. An unlabelled line renders the first sentence of its gloss, which was written as depth, not as a headline, or its humanized id when the gloss is empty; the older advice to leave an obvious deductive connector unlabelled is withdrawn for strengthed lines.

    Objection lines speak in the objector’s voice (“the worry is just chatbots, so no ASI is in prospect”), response lines in the answer’s, and a battery that marks its voice keeps the marker (“the hope: …”). Evidence labels crop at about 56 characters in the graph (lint W5) and must fit their plate at the node’s size tier (TRANSLATION_NOTES L14, pinned by label-crop.test.ts), so distill: where a plain clause cannot fit, keep the subject and the verb and let the gloss carry the rest.

  3. The gloss is the depth tier: the full reasoning, qualifications, asides, source voice, quotes. Glosses are never length-linted.

The three-job test for evidence gloss text, from the corpus survey that motivated evidence labels (AUTHORING_NOTES 2026-06-12): gloss content is either (a) a role tag (“undercut of …”), which is derivable from topology and should be deleted; (b) the warrant, which belongs in the label; or (c) format-meta commentary, which belongs in a # comment. What remains after the test is the genuine depth tier. Job order matters even inside a gloss: put the substantive point first, because displays crop from the end.

Two further conventions from the accessibility passes:

  1. Plain-first, technical-nested: write the gloss in plain language; move a technical restatement to a folded continuation line beginning “technical reading: …”.
  2. Rubric provenance is not reader content. Elicitation citations (“R-STEP S2: …”) go in a trailing # comment on the node line, not in the gloss. Reader-valuable quotes and footnote refs stay in the gloss.
  3. In multi-speaker maps, prefix evidence labels with a speaker tag (“A:”, “L:”, “AL:”); IDs are invisible at graph junctions, so the label carries attribution.

7. Structural idioms

The patterns below carry most of the flagship map (experiments/llm-extraction/iabied-comprehensive-en.argmap; line numbers are as of 2026-07-25 and may drift, so each entry also names the anchor to search for). Excerpts are trimmed; open the real file for the full context. All excerpts are fragments, not standalone files.

The same catalog appears in compressed form in the skill’s ## Structural idioms section, numbered to match these subsections (7.1 = idiom 1, and so on). For a pattern in a complete small file rather than as a fragment, examples/README.md maps each example to the idioms it demonstrates.

7.1 The objection/response triple

The workhorse. An objection statement (the hope or doubt), an objection evidence concluding against the target, and a response undercutting the objection evidence:

# fragment - not standalone (flagship ~line 628, anchor "@c11-readthoughts")
@c11-readthoughts [We'll read the AI's thoughts and catch bad plans] 0.15?:
$c11-readthoughts-obj 0.2? ~@wont-solve-in-time | @c11-readthoughts:
$c11-readthoughts-resp [punishing visible bad thoughts hides them] 0.85? @wont-solve-in-time | @steering-finds-subversion AND $c11-readthoughts-obj: training against legible bad thoughts selects for concealment, not for good ones

FAQ-shaped sources map one-to-one onto rows of these. The -obj/-resp ID suffixes are a mnemonic convention, not syntax. Each line takes its own kind (4.1): the objection line’s is usually the hope’s or the critic’s (hope, testimony), the response’s the book’s own reason (mechanism, deductive, or analogy where it answers with a case), and the undercut takes the kind of its own step, never the objection’s. A gloss-less objection line in a battery has no words of its own to classify; it takes the kind of the objection’s ground, read off the premise statement’s words (the rule the flagship’s classification settled on, 2026-09-08).

7.2 The undercut ladder, including second-order undercuts

Rebuttal, undercut, response, and undercut-of-undercut are one schema applied repeatedly (flagship ~line 283, anchor “$uc-counting”):

# fragment - not standalone
$uc-counting [this argument form fails in ML contexts] 0.3? ~@fragile | @nn-generalize AND @sgd-bias AND $fragile-count:
$uc-uc-counting [the ML rescue may not transfer to alignment] 0.7? @fragile | @gen-not-values AND $uc-counting:

The second line reinstates @fragile exactly to the extent the first line’s rescue fails.

7.3 Linked and convergent, side by side

One hub with both shapes (flagship ~line 269, anchor “$fragile-ev”):

# fragment - not standalone
$fragile-ev 0.9? @fragile | @orth AND @contingent AND @fragility:
$fragile-count 0.7? @fragile | @counting: lottery-ticket prior over goal-space

The first is a linked three-conjunct rule (all needed); the second is an independent convergent sibling on the same conclusion. The test for linked: neither conjunct alone suffices. The flagship’s cleanest statement of that test (anchor “$spread-ev”): “neither end alone shows disagreement; together they are the spread”.

7.4 Convergent siblings instead of a false AND

When a source presents overdetermined routes (“any one of these suffices”), write separate evidences, not one conjunction (flagship ~line 170, anchor “$adv-speed-ev”):

# fragment - not standalone
$adv-speed-ev [speed alone breaks the human range] 0.9? @ai-advantages | @adv-speed:
$adv-copy-ev [copyability alone breaks the human range] 0.9? @ai-advantages | @adv-copy:
$adv-selfimp-ev [self-improvement alone breaks the human range] 0.75? @ai-advantages | @adv-selfimp:

The flagship originally had these as a four-way AND; the repair note in the file records why that was wrong (the book is explicit that no single advantage is necessary).

7.5 Coarse summary plus refinement

The whole book’s case is one coarse line whose refinement holds everything (flagship ~line 261, anchor “$link “):

# fragment - not standalone
$link [the book's claim as one coarse implication] 0.93? @everyone-dies | @if-built: unfolds below into the full case

The coarse strength on a refined line is not a solver input (the refinement replaces it); it is the evidence-side check, displayed against what the refinement delivers. Recommended practice: author the coarse strength as your holistic judgment of the whole implication before trusting the steps; the comparison is a free audit.

What the refinement delivers, and why it sits lower than you expect. The folded line shows the composition of the block under it: the steps composed the way the block wires them. Steps the argument all requires multiply. So a refinement of four required steps at 0.9 each shows 0.656 on the fold, and one of two steps at 0.9 shows 0.810 (both measured 2026-09-12 on a chain standing on its own). That is the price of writing the steps down, and the first time you see it the number looks like a bug. It is the arithmetic of a chain: every step is one more thing that has to hold. A conclusion the rest of the map argues against reads lower still, because the fold is read on the solved joint like every other number.

There is a second factor, and it pushes the same way. The folded reading asks whether the steps are in force and whether what they need holds, given the premises the coarse line itself carries. A refinement usually reaches for grounds the coarse line does not carry (the interior roots its steps stand on), and it pays for each of them. A fold whose refinement is written on firm ground and short chains sits close to its coarse number; a fold four steps deep over five borrowed grounds sits far below it. Both are honest readings of what you wrote.

The rule, stated positively: pick the coarse strength at or below the weakest step you are about to write under it. A chain of required steps can never come out above its weakest step, whatever you believe about how the steps hang together: each step is one more condition on the same event, so the conjunction is at most as likely as the least likely of them. This is arithmetic and there is no way to author around it. If the summary you want to write is firmer than any step under it, put the summary where a summary belongs: a # check: on the conclusion, which records your holistic judgment and lets the solve be audited against it, with the line’s own strength set at or below the weakest step.

Three honest repairs when a coarse number sits above its weakest step, and the source decides between them. The step may be under-elicited, in which case go back to the passage and read its register again. The summary may be your holistic judgment rather than a claim about this chain, in which case it is a # check:. Or, most often, the source gives several grounds and the map wired them as one chain: write them as convergent lines into the conclusion (7.4), and the fold saturates instead of multiplying.

A family of parallel objections that all fail for one reason takes that reason as a premise. The same reading has a consequence for the other direction. When a refinement holds k objections that the source answers with one move, the objections OR upward: each one is another way for the block to deliver, so the folded number climbs above what you meant the family to carry. The repair is 7.9’s shared latent conjunct: name the one reason they all fail as a statement, and conjoin it into every member of the family, so the members stand or fall together the way the source says they do. The flagship’s $hope-harmless is the worked example and the one left undone: four hopes that a superintelligence’s goals never touch us (a digital realm, a cosmos elsewhere, no evolved greed, boredom), and the book answers all four with one fact, that material resources serve almost any goal. The coarse line carries that fact as its premise; the four objections under it do not, so the fold reads 0.268 against an authored 0.07 (measured 2026-09-24 under the one draw; 0.313 on 2026-09-12, before it). Judge the repair by the conclusion’s own value rather than by the fold: a conjunct that repeats the coarse line’s own premise cannot show up in the folded number, because the fold is already read given that premise.

The check. argmap-query fold-audit lists every strengthed refined line in a file with its authored strength, the weakest required step of its refinement, the premises the refinement introduces that the coarse line does not carry, and a flag where the authored strength is above that weakest step:

node mvp/packages/parser/bin/argmap-query.mjs my-map.argmap fold-audit

It is advice and never a diagnostic: a summary written in the author’s own holistic register is a legitimate thing to write down, and the audit only tells you that it is what you wrote.

A layout-driven special case is the coarse hull (see the spine test, 8): when the fine conjunction mixes one cross-region premise with hubs that belong inside the region’s fold, condition the coarse line on just the cross-region premise. The refinement holds the fine line and the local clusters; the collapsed view keeps the cross-region edge.

7.6 The complementary partition (“even if”)

Two routes that would overlap are made disjoint by conjoining the negation of the other route (flagship ~line 638, anchor “$mwb-time”):

# fragment - not standalone
$mwb-hard [the hardness route] 0.9? @misaligned-when-built | @align-hard:
$mwb-time [the timing route, in the solvable worlds] 0.9? @misaligned-when-built | @wont-solve-in-time AND ~@align-hard:

The ~@align-hard conjunct is the source’s own “even if alignment were solvable” made explicit; without it the two routes double-count.

7.7 The balancing evidence

A conditional is vacuous outside its slab, so a map whose every evidence on @c conditions on @a says nothing about the ~@a worlds. If the source asserts the converse, name it (flagship ~line 254, anchor “$no-doom-otherwise”):

# fragment - not standalone
$no-doom-otherwise [if nobody builds it, it kills nobody] 0.9? ~@everyone-dies | ~@if-built: the book's own converse of the title conditional

This is the converse of 4.2: it speaks to the worlds where the premise fails, so it moves the conclusion there and leaves the premise alone, and it keeps a contested base-rate claim on the map, right of a given bar, instead of hiding it in a prior.

Close the set on a spine line (D168, 2026-09-27). A conjunctive line into a claim the map leads with has one failure branch per premise, and the solve fills every branch nothing speaks to at one half, which then shows as part of the headline. So give each premise its inverse where one holds: by the meaning of the two claims (# kind: formal, 0.99, as for the pillar inverses), or in the source’s own words (the sentence’s rubric row and kind). Where neither holds, invent nothing: the fill stays, and the band beside the number shows its share. The flagship’s $link-fine conjoins @if-built, @misaligned-when-built and @mis-ext; $no-doom-otherwise covered the first branch, and the other two now read:

# fragment - not standalone
$aligned-no-extinction [an ASI aimed at what we want does not kill everyone] 0.9? ~@everyone-dies | ~@misaligned-when-built:  # S2, chapter 12's flat "would"; kind: deductive
$harmless-no-extinction [if a misaligned ASI is not lethal, not everyone dies] 0.99? ~@everyone-dies | @misaligned-when-built AND ~@mis-ext:  # by meaning; kind: formal

Keep a by-meaning line strictly by meaning: the second line’s @misaligned-when-built conjunct keeps it off the aligned worlds, where ~@mis-ext says nothing about what an aligned ASI does (without the conjunct the misalignment drag reads 0.032 in place of 0.044, the extra drop being no argument of the book’s). Find the shape with the drag a reader is likeliest to make: before these lines, zeroing @misaligned-when-built left the title claim at 0.427, nearly all of it fill, and about 9 of the 0.708 points at rest were fill too; after them the title reads 0.617 at rest and 0.044 under that drag (measured 2026-09-27, AUTHORING_NOTES of that date; 0.603 at rest and still 0.044 under the drag since the objection rewiring of 2026-09-28, D170).

7.8 Conditioning on an inference (the designed W1)

A policy conclusion that hangs on an implication, not on a fact (flagship ~line 770, anchor “$shutdown-ev”):

# fragment - not standalone
$shutdown-ev [if built means everyone dies, no one may build] 0.93? @shutdown | $link:

Conditioning on @everyone-dies instead would be subtly wrong: doom that were unconditional would justify no ban. This is the rare positive evidence-as-premise; the lint fires W1 by design, and the file says so in a comment.

7.9 The shared latent conjunct (one doubt, many hopes)

When k objections are expressions of one underlying doubt, name the doubt as a statement and conjoin it into every member; elicit its prior once, family-holistically (flagship ~line 435 onward, anchor “$hope-care”; the shared conjunct is ~@no-right-care):

# fragment - not standalone
$obj-cheap 0.7? ~@not-preserved | @cheap-keep AND ~@no-right-care: a sliver of care plus a negligible bill would get paid
$resp-cheap [it would need a reason to pay ours] 0.9? @not-preserved | @needs-motive AND $obj-cheap:

Prefer as the latent the statement the support side already denies, so attack and support quantify over the same worlds. This was the only structure (of six probed) that stayed stable as hopes were added.

7.10 The epistemic-fact reification (norms and credal thresholds)

“A risk no one can bound justifies a ban” is a threshold argument over a credence, which the solver cannot represent directly (facts about credences are not world facts). The pattern: reify the evidence-state as a first-order statement, state the norm as its own statement, and let a near-deductive step combine them (flagship ~line 793, anchor “@risk-unbounded”):

# fragment - not standalone
@risk-unbounded [No one can currently bound the extinction risk below the actionable threshold]: a fact about what has been demonstrated, not about anyone's opinion
@no-gamble [Running a risk no one can bound below the threshold is impermissible] 0.93?: the norm, stated where it can be attacked
$shutdown-fine [unbounded risk + the norm license "don't build"] 0.9? @shutdown | @risk-unbounded AND @no-gamble:
$bounded-escape [a demonstrated bound would dissolve the case] 0.9? ~@shutdown | ~@risk-unbounded:

Note $bounded-escape: the author naming the condition under which their own conclusion lapses. A self-declared off-ramp is both honest and persuasive. The wrong shapes (conditioning on the chance itself, an OR over world-types) are documented as fixture examples/edge-cases/e15-reified-chance.argmap.

7.11 Rebuttal guards (which epistemic state does the defeat presuppose?)

Seven flagship responses only work while the risk is undemonstrated, so each carries the guard conjunct @risk-unbounded AND ... (flagship ~line 688, comment anchor “risk-conditional rebuttal guards”). Structure only, no new numbers: in worlds where the risk is demonstrated bounded, the defeats lapse and the objections revive. Ask this of every response you author (checklist item 8).

7.12 Exclusive alternatives and authored abduction

Rival explanations that cannot both hold are a premise-less constraint factor plus an authored abductive step (examples/09-exclusive-causes.argmap, the pattern catalog):

# fragment - not standalone
$who [an eaten cake means one of the two ate it] 1.0 @alice OR @bob | @cake: the abductive step, stated as a contestable rule
$notboth [they would not both have eaten it] 1.0 ~@alice OR ~@bob: premise-less unconditional constraint factor

The lesson recorded there: abduction is authored, not free. Pinning the effect gives the causes no diagnostic lift by itself; “it must have been one of them” is a premise, and making it a visible, attackable node is the point.

7.13 The direct-assertion pattern (spoken and debate sources)

A flat spoken assertion with no stated grounds becomes an attributed premise-less evidence (experiments/llm-extraction/debate-tang-shapira.argmap):

# fragment - not standalone
$a-blur [A: attention is a blur of what causes what] 0.9? @opaque: the quadratic self-attention transformer "literally is a blur of what causes what" [^t005008]

Sixteen of these carried the debate map. Related: a refusal to give a number (“P(doom) is not assignable”) needs no special syntax; leave the marginal blank and, if the refusal is itself argued, map that argument as an undercut cluster against assignability. How several such lines on one statement combine (pool into one number, or stack, each adding its own weight) is decided (D166 item 4: a line with no premise is a line whose premise is every case, so same-side lines stack where nothing opposes them) and applied after the release: until then the solve pools them, and the lint’s I4 note on every premise-less strengthed line says so. A line that reports one source names the source as a premise instead.

7.14 The parable at zero depth

Narrative belongs in folded gloss continuation lines, not in nodes (flagship ~line 165, anchor “a parable”): a whole illustrative story attaches under one statement, costs no graph structure, and folds away. Use this for the source’s most persuasive prose, which is usually exactly the material that does not decompose into premises.

8. Writing large maps

The flagship map holds roughly 200 statements and 270 evidences at depth 5. The disciplines that made that possible (AUTHORING_NOTES 2026-06-18 onward):

  1. Width, not depth. Every new objection cluster is a sibling under its target, never a deeper chain. When a sub-debate wants an eighth level, promote the deep node to a shared top-level node instead. Node count can triple while max depth stays flat. Width is for distinct considerations (2026-09-25): a sibling line that restates a consideration the map already carries is a second dock of it, and two same-side lines stack, so it is counted twice. Give its passage a second > quote on the existing line instead (10.4 item 18).
  2. A manifest comment block at the top, past about 150 nodes: the coarse spine drawn in ASCII, every shared node listed with its home region and consumers, and the region-prefix scheme stated (@c5-trade, $c5-trade-obj, $c5-trade-resp). Build in dependency order; lint after every region.
  3. Reuse shared grounds aggressively, and annotate each reuse site with a comment naming the home region; otherwise later editors bury cross-references. A handful of high-traffic shared nodes is what keeps maximal coverage finite.
  4. Home shared nodes above the clusters that use them. An evidence folded inside a refinement contributes no edges while folded, so the spine and shared grounds must not be buried inside clusters. A node declared inside a cluster whose every edge leaves it is a “stranded node” (validator W6); re-home it with its consumer.
  5. Folding is a source-structure decision. Where you nest is where readers’ fold boundaries are. Author clusters as refinements under their target; keep shared material outside.

Nesting discipline

Nesting looks like an art but is mostly a mechanical review pass. Real arguments cluster on their own; first drafts nevertheless come out flat (checklist item 5), typically with the clusters already visible as # section-heading comments. Treat that as the diagnostic: section headings are nesting debt. A divider comment organizes the text file; only indentation organizes the reader’s view. If you felt the need for a # ---- divider, the argument just told you where a fold boundary is.

Three tests turn the debt into structure:

  1. Fold-unit test. Would a reader want to collapse this sub-debate to one line? Then give it a wrapper evidence whose refinement holds the cluster (7.9’s hope-battery shape), and author the wrapper’s coarse strength as your holistic judgment of the cluster’s net force; the refinement-vs-coarse comparison then audits you for free (7.5). Smaller version: a statement’s grounds and their evidences nest under the statement.
  2. Burial test. Anything referenced from outside the cluster moves up out of it. A shared ground homed inside one cluster still works, but it renders as a cross-reference burial and, in the worst case, a stranded node (W6). Home shared nodes above every cluster that uses them.
  3. Spine test. The collapsed view must already show the argument’s shape: a folded evidence contributes no edges, so a buried spine disappears from it, and the validator says so (W21, with the linking evidences to lift named in the message). But do not over-correct into lifting every sub-conclusion to top level: that trades a wall of disconnected cards for a crowded one (the He extraction did both in one day: first zero top-level evidences, then fifty top-level cards). Author the top tier deliberately, and keep it coarse: the headline, the sinks, the major route hubs, and the shared grounds the burial test already forces up, roughly 15-30 cards on a large map; every statement hub consumed only within its own region lives one fold down, inside that region. If the source draws its own overview map (a section-2 diagram, an abstract’s roadmap), the flat view should be that overview.

    The mechanics rest on a folding asymmetry: a statement block folds to nothing, an evidence refinement folds to a visible coarse line with its edges intact (7.5). So an edge between two top-level statements must never sink into a statement block. When all its premises are top-level, the evidence simply stays top-level. When it mixes one cross-region premise with region-local hubs (the shape that otherwise forces those hubs to stay top-level and crowds the tier), write it as a coarse hull: a coarse line conditioning on just the cross-region premise, with the fine conjunction and the local clusters in its refinement ($takeover-ev @takeover-doom | @unaligned-asi in the He map, fine five-way conjunction one level down). The spine edge stays visible collapsed, the detail unfolds in place, and the solve runs on the fine line while the hull’s own number becomes the composition of its steps (the folded readout, 10.3), whose gap to the authored coarse strength is then an audit of the summary, never an error to tune away.

    Quick checks: grep -c '^\$' returning zero on a multi-statement map means no spine at all (W21 fires); a top rank past ~40 cards means the tier is set too fine (nothing fires; this one is on you). The ~40 bound assumes one argument. If the map is an atlas, several blocks with :: groups organizing them, read the bound per group: a table of contents is wide on purpose, and nest-audit says so rather than calling it a crowd. One or two free-standing exhibit nodes beside a visible spine are fine (W21 stays silent then); fifteen are not a view, they are a deck of unshuffled cards.

    Run argmap-query nest-audit once the skeleton stands, and again after any restructuring pass: it counts the top tier for you, names the boxes whose opened view is a wide and deep wall, and lists the statements whose support cone is ready to fold, each with the edges that block the fold (reference in mvp/README.md). It counts a cone by what still stands at the statement’s own tier, so once you fold part of a cone under one of its own members the suggestion goes quiet instead of repeating itself. It is advice, not a check: nothing it prints is a diagnostic, and declining a fold it proposes is a normal outcome. Two siblings run the same way. argmap-query group-audit names boxes where three or more objection families stand without a ::group around them and prints the group header to paste, and it flags a declared group whose lines tell no one story (a possible catch-all). argmap-query cook-audit names structure that would steer the solve rather than report it: objections buried two refinement levels below the claim’s supports, and near-duplicate supports stacked on one claim. Both are advice in the same sense, and the call stays yours.

Sometimes all three tests fail and the heading is still real. That happens when the section is a topic, not a fold unit: two arguments that share a file but not a single premise, or a shelf of background facts. Nesting them under a wrapper evidence would be a lie: there is no inference there to summarize. That is what a declared group is for (3.11): write ::id [Label] and indent them under it. The heading stops being debt and becomes a checked, drawn box, and the validator will tell you if the topics you claim are separate have quietly grown a shared premise.

A useful smell figure: the flagship map holds 431 nodes at depth 5. A hundred-node map at depth 1 is under-nested even if every individual line is well-formed; its reader meets a wall of top-level nodes and the fold control does nothing.

9. Limitations

What the format and semantics currently cannot express, with the standing workarounds. None of these block parsing or display; they bound what a solve can mean.

  1. Scope conditionals (FORMAT_DESIGN Q8). “Aligned in the current regime, degrades at superhuman scale” has no first-class form. Marginals capture partial truth, not the conditioning scope; nesting is a partial workaround whose limits fixture e06 documents. The intended direction (D14) is partitioning statements into substatements; undesigned. Until then, statement granularity is the author’s burden.
  2. Undercut fan-out. An undercut names one target. Class-level methodological objections (“this is all unfalsifiable”) attack a family of inferences and end up structurally under-stated as one representative undercut. Mitigations: give the family a shared gate premise and rebut that once; or make the objection a shared Tier-1 ground feeding several undercuts.
  3. No statement re-opening (D33). You cannot declare a statement and attach its refinement later in the file; refinement is physical indentation. Workaround: declare nodes at their refinement site and forward-reference them (IDs are document-global).
  4. Statement-level provenance. ? marks numbers as estimated, but who asserts a claim has no in-format home beyond footnotes and ID prefixes. In multi-source maps this makes scope policing (“does this node belong to this map’s source?”) a manual discipline.
  5. Binary statements only (S8). Categorical or continuous claims must enter through threshold-gate statements (“X exceeds T”).
  6. Credal links are second-order (S12). Nothing computes “if the probability of X exceeds t then Y”; the epistemic-fact reification (7.10) plus a # gate: comment audit is the pattern.
  7. Facts, norms, and future scenarios mix by convention only (S13). The flagship keeps policy conclusions as “should” statements and has no “will X happen” node whose truth would feed back onto its own antecedents. If you add scenario nodes, index them explicitly or the map becomes self-referential.
  8. Independence is assumed (P2) and dependence must be authored (4.4). There is no correlation annotation.
  9. Solver cost grows with treewidth (P5). The corpus solves in seconds at treewidth about 7 to 9; a much more entangled map may not. Width-not-depth authoring also keeps treewidth down.
  10. Comment-layer slots are conventions. # check: and # gate: are invisible to tools other than the solver readouts, and nothing validates them structurally.
  11. Cross-map ID reuse is unchecked. Reusing an ID across maps is string coincidence; verify the propositions match before treating them as the same claim (a debate’s “prepare an off-button” is weaker than the book’s @shutdown).
  12. Authoring cost is real (RISKS §2). Mapping is slower than prose. The mitigations that exist today are LLM extraction with human steering, and the rubric discipline that keeps the numbers honest (RISKS §4); neither removes the labor, they redistribute it toward review.

10. Checking your map

Three mechanical gates, in order: lint, parse, solve.

10.1 The lint

python3 tools/argmap-lint.py path/to/your.argmap

Zero errors is mandatory. The codes (full table in tools/README.md):

Code Meaning Author action
E1 duplicate ID (one namespace across @/$) rename
E2 dangling reference fix the ID
E3 ~$id rewrite as an undercut (3.6)
E4 probability outside [0,1] fix
E5 v0.3 pair without argmap-version: 0.3 declare the version
E6 malformed pair (0.9/, /0.2) write both members
E7 ::id in an expression (3.11) a group takes no part in inference; reference a member
E8 a probability on a :: line (3.11) groups have no credence slot; delete the number
E9 > outside an annotation block (3.12) move the quote under its node, before that node’s first child
E10 > with no quote text write the quote or delete the line
W1 evidence-in-premise, not undercut-shaped usually a polarity slip; legitimate only for deliberate conditioning-on-an-inference (7.8), then say so in a comment
W2 directed cycle usually fine (mutual rebuttal); check it is not a zero-negation support cycle
W3 prose line resembling a node you lost a sigil; fix it
W4 footnote used/defined mismatch fix
W5 evidence label past ~56 chars distill the warrant; depth to the gloss
W7 pair sums > 1 declared two-sided conflict or infeasible residual; confirm intended
W8 pair 0/0 drop it
W9 pair entangled with undercut shape check what the opposed side actually asserts
W11 authored 0 strength you probably mean an unstrengthed line
W15 quote line with no [^locator] (3.12) add the locator; provenance is the point
W16 leftover [^ inside quote text (3.12) only a trailing ref is the locator; fix the stray or doubled one
W17 ` # ` inside quote text (3.12) quote lines have no trailing comment; move the note to a #[…] line
W18 quotes before the end of the gloss prose (3.12) reorder: gloss first, then quotes
W19 > under a declared version below 0.3 declare argmap-version: 0.3
W20 retired ~"…" still in a gloss (3.12) migrate it: > line, plain marks, or the echo pattern
W21 buried spine: top-level statements unconnected in the flat projection (8) lift the linking evidences it names to the top tier
W22 zero-node file the parse went wrong; read the counts, fix the file
W23 a bare point and a # check: on one head line (4.3, 4.7) drop one, or write the unargued remainder as a pair beside the check
W24 an @id/$id token in prose that resolves to nothing fix the typo, or drop the sigil if it is not a pointer
W25 a bare point beside mapped support (4.7) derive: move the number to a # check:, or write the remainder as a pair
W26 malformed # check: token (3.7) a point in [0, 1] or lo..hi with lo <= hi
I1 stats; isolated statements connect or delete isolates
I2 block inventory (multi-block files) read it; confirm the split you intended
I3 unstrengthed-line inventory (4.7) the lines that compile inert; commit or drop each before calling the map done
I4 a premise-less strengthed line: the reading note (7.13) the solve pools it with the other numbers on its conclusion for now; the stacking reading is decided (D166 item 4) and lands in a later release; a line reporting one source names the source as a premise
I5 parallel leaves: two or more sibling lines with one conclusion and one premise set (settled by D166) they are same-side shares of one population, independent where nothing opposes them; one argument written twice merges
I6 coinciding pairs from different premise sets into one statement (settled by D166) where their premises hold together they are one draw and the strongest share on each side holds (4.4), so agreeing lines read as either one alone

Two caveats: the stranded-node check (W6) lives only in the TypeScript validator (visible in the editor), not in this lint; and on expression-valued conclusions (@a OR @b left of the given bar) the lint’s undercut-shape family (W1/W9 and same-slab W7) deliberately stays single-ref, so near-misses there are the editor validator’s job (see tools/README.md, update 2026-07-25).

Warnings are advisory and some are load markers on purpose: the flagship map ships with two deliberate W1s. The discipline is not “zero warnings”; it is “every warning has an explanation you could put in a comment”.

10.2 The parser

Open the file in the editor (or run the parser test suite if you work in the repo). The editor shows diagnostics inline, including the validator-only warnings (W6 stranded node, W10 version gate). Without the editor (a standalone tool bundle), a successful solve_map.py run doubles as the parse gate: it loads the file through the real parser.

10.3 The solve

cd experiments/solver-prototypes
python3 solve_map.py path/to/your.argmap --top 10

Read the four sections: statement gaps (authored or check value vs solved), evidence tensions (authored strength vs achieved), spectator gaps (a folded line’s authored strength vs the conditional the whole network delivers through its refinement, a bench readout) and composition gaps (the same authored strength vs the composition of the refinement’s own steps, which is the number the editor shows on the folded line). Then query the nodes you care about:

python3 solve_map.py path/to/your.argmap @headline '$main-step'

(The forced interval is a bench instrument of the retired uniform reference: add --reference d36 --band. It needs the optional band_probe.py next to solve_map.py, and --influence needs influence_probe.py; the plain readout needs neither.)

Interpreting what you see:

  1. A large statement gap: the mapped argument does not deliver the authored or checked belief. Revise structure or strengths if the argument is misstated; add named missing evidence if real support is unmapped; otherwise keep the badge, it is a finding.
  2. A large evidence tension: the map holds the line below the strength its author wrote (under D161 only that direction tints; a line carried above its strength is the fill); look for an overlooked conflict with neighboring lines.
  3. A composition gap: the refinement’s own steps, with the claims they rest on that nothing inside argues, compose to something different from the coarse strength its author wrote on the folded line (“steps outrun summaries”, or the reverse at the spine). That composition is the folded line’s readout in the editor (since 2026-09-08, D161 item 9), marked by a dotted track and, since 2026-09-12, tinted like any other gap (D38 as amended); decide which side is wrong, both states occur in practice. The spectator gap beside it is a bench readout only: the network’s delivered conditional also reads the conclusion’s other parents, so it is not the line’s own number and not a verdict on either side while SOLVER_SEMANTICS S21 is open.
  4. Numbers never move to make badges disappear (5.2.2). Structure moves, named evidence is added, or the badge stays and means something.

The solver needs numpy, scipy, and node. If it is unavailable, lint plus parse still validate everything structural.

For translated maps, python3 tools/translation-parity.py BASE TR verifies the translation touches only free-text spans.

10.4 The pitfalls checklist (for an adversarial auditor)

Work this list against a finished map as if you were paid to find the double count. Each item says what to look for, how to see it (the lint code or argmap-query audit that catches it, or the reading that does when no tool can), and the fix. The tools are python3 tools/argmap-lint.py FILE and node mvp/packages/parser/bin/argmap-query.mjs FILE <audit>; every audit is advisory and prints the shape, and the call stays yours. The rationale for each item is 4.7.

  1. A pin beside mapped support. A statement with an authored point and at least one strengthed line concluding into it; the point counts that line’s consideration twice (T1). See it: W25. Fix: derive it (the number moves to # check: at its register’s interval), or write the unargued remainder as a pair.
  2. A pin and a check on one head line. The point pins the statement to itself while the check silently disagrees; the solve reports a tiny gap that reads as convergence. See it: W23. Fix: drop one; on an interior statement the check stays and the point goes, or becomes a residual pair if the text licenses one.
  3. A pinned root homed in a box it does not feed. A pinned root declared inside a refinement box or ::group, consumed only outside it, wired to nothing inside it: shared ground parked where the map holds its likeliest antecedents. See it: isolate-audit (each row names the container and the outside consumers). Fix: derive it from the container if the inference test passes, re-home it with a consumer, or say in the gloss why it stays.
  4. A restatement wired as inference. A line whose premise and conclusion are the same proposition or episode at two docks (the @psychosis / @c13ws-retrain shape): one consideration counted twice through a line that carries no inference. See it: restate-audit lists lines whose premise and conclusion cite only the same locator; the reading is “does the passage make this step, or say the same thing twice?” Fix: delete the line and cross-reference the docks in a comment, or merge them into one statement.
  5. Two lines sharing a premise into one conclusion. Same-polarity lines into one statement whose premise sets overlap, combined as if independent. See it: shared-cause, one row per shared statement. Fix: either the overlap is the deliberate world-layer device (the shared conjunct, 7.9) and a comment says so, or factor the shared span out, or merge the lines.
  6. A statement AND-ed with its own derivative. A junction @x AND @y where some line derives @y from @x (the flagship’s $wst-race): the consideration enters the junction twice, and no static audit sees it. See it: the reading. For every AND, trace each premise’s ancestry two hops up (argmap-query neighbors ID --hops 2); a premise that appears in another premise’s ancestry is the hit. The covariance probe (experiments/solver-prototypes/covariance_probe.py) finds it at scale as premises that co-move under one root’s ablation. Fix: drop the duplicate premise, or re-elicit the line conditional on the derivative alone.
  7. A floor elicited from the total. A derived statement carrying a residual pair whose floor sits at, or near, the register’s point: the pin again with extra steps (T3’s hazard; the spoke’s variant C). See it: the reading. A pair beside a check is lint-silent by design, so compare the two: a floor that would close the badge on its own is the total, and a floor with no passage behind it in the gloss is the total by default. Fix: empty the residual unless the text names a second unwired ground (floor at that ground’s register, passage in the gloss) or a remainder (“to name a few”: 0.2?/0?).
  8. A posterior elicited as direct evidence. A root’s pair read off a register that already discounts what the author believes downstream, so the solve applies modus tollens a second time. See it: the reading; the signature is an author hedging a root because of a conclusion the map also derives from it. Fix: there is no exact correction. Keep the table’s widths, say in the trailing comment that the register is a posterior, and present the interval and the band rather than the midpoint.
  9. A definition as a p = 1 statement. A terminology node (@asi-def [ASI means ...] 1) conjoined into premises: it solves 0.988 under the ridge and taxes every junction it joins by 1/400 (T7). See it: grep the marginals for ` 1: and 1?:`; the lint is silent (0/1 is a smell without a code). Fix: terminology goes in a gloss or a glossary surface; a definition that does inferential work becomes p = 1 evidence lines, spelled as the converse pair (T6).
  10. A malformed check interval. # check: holding a token that is neither a point in [0, 1] nor lo..hi with lo <= hi (reversed bounds, a bound past 1, a single dot, a dangling ..). The check is display-only, so nothing fails downstream; the badge simply disappears while the comment reads as if it counted. See it: W26. Fix: write the point or the interval.
  11. An unstrengthed line left as structure. A line with no strength compiles inert: the relation is drawn and nothing is asserted, so a map that “maps” an argument through it says nothing about it. See it: I3, the inventory. Fix: commit a strength by rubric, or delete the line if the relation was never meant to carry.
  12. A label whose pronoun has no antecedent. “It must hold across the gap between weak before and lethal after” (the flagship’s @before-after-gap before 2026-08-23) reads cleanly in source order and as nothing in the outline, the graph card or a search hit. See it: the reading. Read every statement label cold and out of order (sort the outline, or read the argmap-query roots listing); any label that needs the line above to say what “it” is fails. Fix: put the subject in the label (“Alignment must hold across the gap …”).
  13. A bundled label. A label carrying two propositions (“X and Y”), or a claim fused with what follows from it (“X, so dismiss Y”): two variables in one slot, and a reader cannot tell which half a line argues for. See it: the reading; grep labels for ` and , so , therefore`, and ask of each hit whether the two halves could be true separately. Fix: one proposition per statement, or derive the fused conclusion from its halves with one line per ground.
  14. A refusal written as a number. The source says “nobody knows” and the map carries 0.5, or a confident point, or a narrow pair, with no gloss. See it: the reading; a pair wider than 0.30 with no gloss sentence, or a 0.5 with none, is the signature. Fix: 0.1?/0.1? plus the gloss sentence naming the passage, or a blank marginal where the refusal is the author’s own.
  15. A pair whose shape contradicts its register. A categorical the source leaves no room against, written two-sided; an objection the source denies, written 0/p with no floor; a hedged claim written as a flat pair. See it: the reading, against the table in 4.7.4 and the class token in the trailing comment (a pin with no class token is unaudited). Fix: the table’s row, or a per-node pair from the text with the rationale beside it.
  16. A restatement under two labels, and the wrong door. A line whose premise and conclusion describe one event under two labels (@asi-soon “ASI arrives this century” feeding @if-built “anyone builds an ASI” on the flagship until 2026-09-25): the line carries no inference, its gloss tends to carry the source’s real reason (there, the chapter 12 race) that its premises do not name, and a what-if that zeroes the premise leaves the conclusion at the reference fill, which a reader sees as a broken sum (50%). Its companion is a claim whose only consumer is that line (@int-power, chapter 1’s power, reached the title only through the build), so a whole chapter’s case enters the argument through the wrong door. See it: restate-audit only where the docks share a locator; otherwise the reading, which the drag makes cheap: for each strengthed line into a hub, zero its premise (solve_map.py --override) and read the hub; a landing near 0.5 is the shape. consumers on each top-level statement finds the single-consumer case; inverse-audit --all lists the hubs a check prices and no inverse line does. Fix: give each label its own proposition (feasibility here, the build there), put the source’s reason in the premises, wire the orphan where the source uses it, and write the analytic inverse where one holds by the meaning of the two claims (~@if-built | ~@asi-soon, deductive, no passage needed). Rule: AUTHORING_NOTES 2026-09-25.
  17. An evidence label that is not the step. A strengthed line with no label, or one whose label is a warrant fragment, a figure the gloss does not unpack, or its premise’s label said again. The reader meets the line before its premises (section 6 item 2), so none of these tells them what the step is about or where it lands. See it: the reading. List every strengthed line’s head sorted by label, so each is read without its neighbours: grep -oE '^\s*\$[A-Za-z0-9_-]+ (\[[^]]*\] )?[0-9.]+\??' FILE | sed -E 's/^\s+//' | sort -t'[' -k2 (the rows with no bracket sort first: those lines have no label; on the flagship, 363 rows and 82 of them bare before the 2026-09-25 pass, 0 after). Ask of each: does it name its subject, and does it say which conclusion the premises give? Fix: one plain clause, premise to conclusion (“our intelligence remade the planet, so it is powerful”), in the objector’s voice on an objection line. Parallel lines into one claim need different words: cook-audit reads a label overlap of 0.75 or more as a duplicate support. Rule: AUTHORING_NOTES 2026-09-25, the label pass.
  18. One consideration at two docks. Two strengthed lines into one claim that are one argument read twice: one motive under two labels ($ext-core and $ext-incidental into @mis-ext on the flagship until 2026-09-25), a rule beside its instance ($soon-frame beside $built-ev), one answer the source gives in two paragraphs (the hired hands and the robot bodies into @wed-lose), or a premise that restates the conclusion ($fragile-count). Same-side lines are independent draws that stack where nothing opposes them (D161, D166), so the second dock counts the consideration twice and the claim reads above what the source’s one argument delivers; at rest a saturated claim hides it, so a drag does not find it either. See it: the reading, and shared-cause is its trigger: every row it prints (lines into one claim sharing a premise, outright or through a refinement) is a pair to read against the source, asking whether the two passages are one answer. cook-audit catches only near-identical labels, and a pair with disjoint premises (a rule and its instance, one answer in two paragraphs) shows on no instrument, so read the lines of every multi-line conclusion side by side against their passages. Fix, by the rule one consideration, one dock: a second passage on the same consideration becomes a second > quote on the same line, never a second line; where a rule feeds an instance, the rule becomes one statement with its own box and one line takes it; where two passages are one answer, the lines merge into one (conjoin the premises, or OR them where either suffices: one line, one coin); where a line sits at the wrong door, re-aim it. Never nest the fix inside an evidence. Rule: AUTHORING_NOTES 2026-09-25, the cleanup.

Run the mechanical half first (lint, then the five audits: isolate-audit, restate-audit, shared-cause, cook-audit, pinned-roots, and inverse-audit --all for the hubs a drag would expose), then the reading half over the statements pinned-roots lists, since those are the rows where items 3, 7, 8, 14 and 15 live, the drag over the hubs for item 16, the sorted label listing for item 17, and a side-by-side reading of every multi-line conclusion for item 18, starting from the rows shared-cause prints.

Appendix A: a complete worked example

The file below is complete and lints clean as shown (zero errors, zero warnings). It exercises: convergent routes, a linked conjunction, a refinement with an evidence-side check, an undercut, a reinstating undercut-of-the-undercut, a rebuttal, check credences, ? discipline, and a footnote.

---
argmap-version: 0.2
title: "Protected bike lanes and cyclist safety"
description: "AUTHORING_TUTORIAL.md Appendix A: worked example."
date: 2026-07-25
---

# Headline first (convention). Derived statement: no authored marginal,
# a check credence instead (residual authoring rule).
@lanes-safer [Protected lanes reduce cyclist injuries per trip]: the headline claim  # check: 0.8

# Route 1: observational. The coarse line refines into the per-trip
# reading; its 0.7? is the evidence-side check against the refinement.
@study-drop [Injury rates fell after protected-lane installation] 0.9?: city-level before/after counts [^lusk]
$obs-route [before/after data carries the claim] 0.7? @lanes-safer | @study-drop: unfolds into the per-trip reading below
  @exposure-ok [The drop is not explained by reduced cycling] 0.8?: ridership rose over the same period, so per-trip risk fell
  $obs-fine [per-trip injuries fell while ridership rose] 0.8? @lanes-safer | @study-drop AND @exposure-ok: linked - both facts are needed for the per-trip reading

# Route 2: mechanism. Convergent sibling of $obs-route (independent
# routes, so separate lines, not an AND). Independence audited: the
# mechanism does not rest on the before/after data.
@separation [Physical separation removes the main collision type] 0.9?: most serious urban cycling injuries involve motor vehicles
$mech-route [the design removes the dominant injury mechanism] 0.75? @lanes-safer | @separation:

# The objection: grants the data, denies the inference from it
# (an undercut of $obs-route, not a rebuttal of the claim).
@confound [Cities add lanes where cycling is already safest] 0.5?: selection: lanes go where streets are calmest
$uc-obs [selection could explain the before/after drop] 0.6? ~@lanes-safer | @confound AND $obs-route:

# The response: an undercut of the undercut (reinstatement). The
# selection story predicts no drop at quasi-random sites.
@natural-exp [Some installations were sited quasi-randomly] 0.7?: construction-driven and court-ordered sitings
$resp-uc [quasi-random sites show the same drop] 0.8? @lanes-safer | @natural-exp AND $uc-obs:

# A rebuttal (attacks the claim itself, so no $-conjunct).
@risk-comp [Riders take more risks when they feel protected] 0.4?: the risk-compensation hypothesis
$rebut [risk compensation could offset the design gain] 0.3? ~@lanes-safer | @risk-comp:

[^lusk]: Lusk et al., "Risk of injury for bicycling on cycle tracks versus in the street," Injury Prevention 17, 2011.

What to notice:

  1. @lanes-safer is derived, so it carries # check: 0.8 and no authored marginal.
  2. Frontier roots (@study-drop, @separation, @confound, …) keep authored values, all ?-marked as estimates.
  3. $uc-obs conditions on $obs-route (undercut); $resp-uc conditions on $uc-obs (reinstatement); $rebut conditions on neither (rebuttal).
  4. The refinement under $obs-route makes its 0.7? a displayed check against what $obs-fine delivers, and $uc-obs re-aims onto the refinement’s delivery line when unfolded.

The actual solver readout for this file (solve_map.py at its default, the network reference G that ships since D161, with the one draw of MATH §3.8; measured 2026-09-23), abridged:

appendix-a.argmap: 7+5 vars, width=6 | 0.0s, conv=True, reference=g, cell-draw=chain
  largest statement gaps (authored/check -> solved):
    @lanes-safer                0.80 -> 0.848  |d|=0.048
    @exposure-ok                0.80 -> 0.799  |d|=0.001
    ...
  largest spectator gaps (authored coarse ~> delivered by refinement):
    $obs-route:delivered-by-refinement    p=0.70 ~> q=0.856 |gap|=0.156
  largest composition gaps (authored coarse ~> in force by composition, the folded readout):
    $obs-route:in-force(composition)      p=0.70 ~> q=0.549 |gap|=0.151
  badges: 0 statements more than 0.10 outside the check interval

Reading it: the mapped argument delivers a little more than the check credence 0.8 on the headline (solved 0.848, inside the 0.10 badge threshold), so the map carries the stated belief; whether the author’s 0.8 was too cautious or the map is missing a qualifier is the question the badge would ask if the gap grew. The rebuttal is what holds the headline there (4.4): where risk compensation holds, $rebut stays in force in 0.25 of those cases, near its authored 0.3, and the headline reads 0.737 there. Under the rule before 2026-09-23, which weighed the routes and the rebuttal as independent evidence, the routes outvoted it (in force in 0.10 of those cases, the headline 0.836 there and 0.890 overall; both measured 2026-09-23 on the solved joint). The sub-0.01 gaps on the roots are the solve meeting each point to its resolution, not tension. The composition row is the number the editor shows on the folded $obs-route (10.3): the refinement’s own steps, with the claim they rest on that nothing inside argues (@exposure-ok, 0.8) and the selection undercut where its answer fails, compose to 0.549, well under the authored 0.7. The coarse strength promises more than its own steps deliver, which is the “decide which side is wrong” case of 10.3: either the 0.7 is too generous, or the exposure claim deserves an argument of its own. The spectator row above it is a bench readout (solve_map.py prints it; the editor’s reader surface does not): the conditional the whole network delivers through the refinement, 0.856 here, which also reads the headline’s other route ($mech-route) and is therefore not the line’s own number (SOLVER_SEMANTICS S21, open). Under the retired uniform reference (--reference d36, the readout this appendix quoted until 2026-09-08) the same file reads @lanes-safer 0.818, the composition 0.520 and the delivered conditional 0.837 (re-measured 2026-09-23; the composition figures this paragraph quoted until then, 0.690 and 0.668, predate the composition’s counting of a route’s unargued interior premises, which landed on 2026-09-08).

Appendix B: cheat sheet

@id [label] p?: gloss                     statement; p optional, ? = estimated
$id [label] s? CONCL | PREM: gloss        evidence; s optional; | optional
$id s ~@x OR ~@y: gloss                   premise-less constraint factor
~@id                                      negation (never ~$id: E3)
AND / OR                                  linked / convergent; parens to mix
::id [label]: gloss                       declared group; no credence, never in an expr
  indented node lines                     under @/$: refinement (replaces parent unfolded)
                                          under ::: membership in the group
  indented prose                          folds into the gloss above
  > verbatim text [^locator]              quote line; one line, no trailing comment
# comment                                 full-line or trailing
#[key: ...]                               annotation comment (per-file free in parity)
# check: p  or  # check: lo..hi           display-only credence (derived stmts); badge = distance to the interval
# kind: word                              on an evidence line: formal, deductive, mechanism, empirical (n=<int>), testimony, analogy, hope
# gate: q($e) >= t => @c                  threshold audit (comment layer)
[^ref] ... [^ref]: source                 footnote citation
---: argmap-version: 0.3                  required for slash pairs (s+/s-, p+/p-) and > lines

Number rules: elicit as “assume the premises; how likely is the conclusion?”; ? on rubric-derived values, bare only for source-stated numbers; derived statements get checks, not pins; no authored 0/1; fix arguments, not numbers, after the first solve.

Counts (D161, 2026-09-08): a claim’s firmness is its width (0.7/0.1 8 flips, 0.85?/0.05? 18, a point 200, the cap); a line’s is its # kind: (formal hard, deductive 1000, mechanism 64, empirical and testimony 16 or a larger stated n=, analogy and hope 4, no key 16). The tint lights only outside what was written (below a strength, outside an interval, off a point, past 0.01); inside is the fill, uncoloured. What-if: your number replaces the author’s on that claim as a point at the cap; on a root the map follows forward, on a conclusion the author’s case retreats where it is softest (free premises, hopes and analogies, judgments, flat assertions, mechanisms, in that order); “reached” on the adjustment row is the refused residual (W4-E renamed it from “held at” on 2026-09-08, so “held at” is the count’s phrase alone; “met at” became “reached” on 2026-09-27).

Undercut schema: $u q ~C | grounds AND $target. Ask: which inference does this objection grant, and which does it deny?

Review checklist, one line each: provenance traced; undercut targets typed; overlaps merged/factored/partitioned or declared; no dangling sub-conclusions; clusters nested, shared grounds top-level; full-source coverage pass; lint + parse + solve smoke; defeat presuppositions guarded; multi-voice overlaps deduplicated.

Check: python3 tools/argmap-lint.py FILE, then cd experiments/solver-prototypes && python3 solve_map.py FILE --top 10.