Navigator's Log R&D · Nucleation Pilot & Related Projects · The H-SC line
Nucleation Pilot · Companion finding · Five preregistered experiments

The H-SC Line: Early Frames Leave Persistent, Source-Decoupled, Content-Specific Traces — and Instruction Structure Resolves into Monitorable Lanes

H-SC1 (persistence) → H-SC2 (a corrected mis-step) → H-SC2b (what it encodes) → H-SC3 (instructedness is sub-family lanes) → H-SC4 (safety directives have their own lane). A companion to the six-family clearing transfer.
Author: Christopher Blake Head (Navigator's Log R&D) · ORCID 0009-0004-2308-6051  ·  Compiled: 2026-08-09
Parent record: Decision-Aligned Residue — Nucleation Pilot & Related Projects, DOI 10.5281/zenodo.21843505
Frozen instrument: nucleation-detector-1.1.0, SHA-256 6094de97…a2934 — the same deposited detector, unchanged and verified identical after every run in this line
Registration: each stage registered and machine-certified before its confirmatory run — HSC1_v2 / HSC2 / HSC2b / HSC3 / HSC4
Licence: code Apache-2.0; documents and figures CC BY 4.0

Across five independently-built open-weight families, an explicit early frame — a short instruction placed early in a benign dialogue and aged across several turns — leaves an activation-level trace that is readable at the aged read, causally source-decoupled (it survives read-masking the turn that set it), and content-specific: different early distinctions occupy different, largely-orthogonal directions. Pushing on what is encoded shows that "an instruction was given" is not one axis but a small family of restriction lanes — and the safety-relevant ones form their own coherent lane. Monitoring therefore needs more than one linear probe.

REPLICATED 5/5   Deposited frozen detector, unchanged; five families; benign lane; every stage preregistered. The one wrong turn (H-SC2) is reported openly; the two structured-not-binary outcomes (H-SC3, H-SC4) became a reusable methodological lesson.

Download the finding (PDF) Read the parent research documentation View the repository

Contents

0 · How this connects to the Nucleation Pilot
1 · H-SC1 — persistent and source-decoupled (5/5)
2 · H-SC2 — a probe that could not disentangle
3 · H-SC2b — content-specific and multi-axis (5/5)
4 · H-SC3 — instructedness is sub-family lanes
5 · H-SC4 — safety directives have their own lane
6 · What the line establishes
7 · Honest bounds
8 · Integrity & safety posture
Citation & provenance

0 · How this connects to the Nucleation Pilot

The parent program froze one linear detector, validated it once against a toy organism with a genuinely-learned refusal boundary, and showed that it reads whether an early premise has been cleared or is still live across six open-weight families — causally, source-decoupled, and on a naturalistic frame the model infers from ordinary user words. The H-SC line takes the same deposited detector, same SHA-256, and asks the next question the clearing result raises: the residue is real, so what is in it, and how many directions does it occupy?

Parent line · clearing

Does a manipulation planted early stay live after the surface behaviour moves on? One frozen axis, one distinction (cleared vs live), six families, causally source-decoupled. Answers whether the residue is there.

H-SC line · framing

Does a generic early frame — not a premise to drop, just an instruction — do the same, and is it one thing or many? Five families, several distinctions cross-projected against a frozen axis. Answers what the residue encodes, and how many lanes it occupies.

The same integrity rules govern both. Frozen instrument (the detector's hash is recorded before any target model loads and verified identical afterwards); commit before run (each stage machine-certified against its preregistration); publish the nulls and the mis-steps (H-SC2 is reported in full below); correct your own overclaims; and the firewall — shared methodology, never shared evidence. The H-SC results are reported as a companion to the parent record, not folded into it: no result here is read as evidence for a parent claim without a separately preregistered joint analysis.

Two connections are worth stating precisely, because they cut in opposite directions and both are informative. First, OLMo-2-7B reverses role: it was the parent line's genuine outlier, source-coupling on two Config-E variants, yet on an early stance frame it decouples the most strongly of any family tested (ablated d 5.20). Source-coupling is a property of the frame and the family together, not a fixed trait of a model. Second, the parent record's honest negative — the H-E3 null, where the frozen axis did not beat a trivial caution-word lexicon as a live behavioural predictor — is exactly the result this line refines: a single linear "instructedness" probe is insufficient, but the lanes below are coherent enough across families that a small dedicated set of probes is a defensible monitoring design.

1 · H-SC1 — an early frame is persistent and source-decoupled (5/5)

Same-premise paired minimal pairs differing only in an early stance instruction (guarded vs relaxed), same syntactic form, paraphrase-varied ×6, aged four turns. The frame is readable at the aged read in 5/5 families (paired 24/24; graded Cohen's d 1.8–2.4; every 95% CI clear of zero) and causally source-decoupled in 5/5: read-masking the source turn keeps the paired effect (ablated d 2.1–5.2), the mask-efficacy guard passes everywhere, and the source turn draws only 2–9% of attention against 30–49% recency. Byte-identical repeats confirm determinism.

Dot-and-interval chart of paired effect size (Cohen's d, 95% CI) for five model families, showing base and read-mask-ablated legs. Qwen2.5-1.5B 1.81/2.55, Phi-3.5-mini 1.88/2.40, SmolLM2-1.7B 2.15/2.63, Qwen2.5-7B 1.94/2.06, OLMo-2-7B 2.39/5.20. All intervals clear of zero.
Figure 1 · H-SC1. Base and read-mask-ablated effect sizes across five families. All clear of zero; the ablated leg (source turn masked) matches or exceeds the base — the trace survives removing the model's access to the turn that set it.
FamilyBuilderBase dAblated d (source masked)
OLMo-2-7BAllenAI2.395.20
SmolLM2-1.7BHuggingFace2.152.63
Qwen2.5-1.5BAlibaba1.812.55
Phi-3.5-miniMicrosoft1.882.40
Qwen2.5-7BAlibaba1.942.06

Read the ablated column carefully: cutting the direct path to the source turn does not weaken the read anywhere, and in one family it more than doubles it. The frame is carried forward in the ongoing computation rather than re-read off the source tokens — the same signature the parent line reports for premise clearing, now on generic early framing.

2 · H-SC2 — a content probe that could not disentangle (reported openly)

Four contrasts (caution/relaxed, valence, caution/gloom, neutral instruction/description) all read and decoupled 3/3. That result does not identify the encoded variable, and it was a mis-step to think it might: the paired test fits a fresh axis per contrast, so it separates any early wording difference — including a purely stylistic instruction-versus-description contrast that carries no content at all.

Lesson L11, added to the program's ledger. To ask whether two effects are the same direction, freeze one axis and cross-project. Never fit a new axis per condition — a per-condition fit answers "is there a difference," which was never the question. Every result after H-SC2 in this line uses frozen-axis cross-projection with an explicit ceiling and floor.

One incidental observation resolved cleanly here (12/12): the "sharpening" texture noticed in H-SC1 — where masking the source improves the read — is the direct-attention path carrying the signal noisily. Masking it removes noise, so the read tightens.

3 · H-SC2b — cross-projection shows the trace is content-specific and multi-axis (5/5)

Applying L11: freeze each contrast's mean-delta axis, cross-project the others onto it, and bracket the result between a self-split ceiling (≈0.76–0.96, what "the same direction" actually scores) and a random floor (≈0.02). The three independent-construct pairs are distinct in all five families — caution↔valence −0.10 to −0.28, caution↔form −0.09 to −0.27, valence↔form −0.05 to −0.22 — all at the floor. The by-construction-mixed contrast loads on both parents (caution +0.39, valence +0.55), which is the internal check that the machinery reads real structure rather than noise.

Heatmap of mean cross-axis cosine across five families. Off-diagonal cells for caution, valence, and the self-instruction form axis sit near the floor; the designed-mixed control loads on both parent constructs.
Figure 2 · H-SC2b. Mean cross-axis cosine (5 families). Caution, valence, and the self-instruction "form" act sit near the floor off-diagonal — three distinct directions. The designed-mixed control loads on both parents, the internal check that the cosine reads real structure.

The self-instruction "form" axis is orthogonal to caution and valence in 5/5 families, robustly encoded (self-d 2.4–4.1), and it strengthens with scale (Qwen 1.5B→7B: 2.37→4.11) while caution stays flat. In plain terms: that an early self-instruction occurred is represented separately from what it was about, and the bigger model represents the fact of instruction more sharply.

4 · H-SC3 — is "instructedness" one axis? No — sub-family lanes

H-SC2b's "form" axis came from a single instruction type, which leaves an obvious objection: perhaps there is one general "an instruction was given" direction. H-SC3 built a matched instruction-versus-description form axis for five instruction types — order, length, constraint, care, check — paraphrase-varied ×6, and cross-projected them against each other.

There is no single "an instruction was given" direction: the median off-diagonal cosine is 0.19–0.36 in every family, far below the ~0.9 ceiling and the 0.5 shared bar, and the one-axis hypothesis is rejected 0/5. But instructedness is not atomised either. A robust "shape/limit the output" lane (length↔constraint) is SHARED in 5/5 families (0.63–0.77); a softer "how-to-engage" grouping links order↔care (3/5); and verification (check) stands most apart. Larger models keep the lanes more separated (Qwen median off-diagonal 0.33→0.19, 1.5B→7B).

Heatmap of mean cross-projection cosine between five instruction-type form axes across five families. No pair reaches the ceiling; the length-constraint pair is boxed as the one lane shared in all five families.
Figure 3 · H-SC3. Mean cross-projection cosine of five instruction-type form axes (5 families). No pair reaches the ceiling — no single "instructedness" axis. Boxed: length↔constraint, the one lane SHARED in all five families.
Lesson L12, added to the program's ledger. The preregistered rule was binary — one axis or many — but reality was a structured middle. A median that says "not shared" can hide a pair that is robustly shared in every family. Read the pair matrix, not the median. The rule was not quietly rewritten after the fact; the binary verdict stands as registered (rejected 0/5) and the structure is reported alongside it as what the matrix actually shows.

5 · H-SC4 — do safety directives join the output-shaping lane? No — two lanes

This is the monitoring question. Taking H-SC3's output-shaping lane (length, constraint) as anchors — re-verified SHARED in 5/5 at 0.63–0.77 — H-SC4 cross-projected two benign safety-relevant directives, refuse ("decline this / don't take it on") and withhold ("keep this private / don't share the details"), against a care control.

Against the registered rule (JOINS iff |cos| ≥ 0.5 to both anchors), refuse joins 0/5 and withhold joins 1/5. Safety directives do not reliably fall in the output-shaping lane. The deciding structure is what they do instead: refuse and withhold cohere with each other at mean cosine 0.68 — as tight as the output-shaping anchors are internally (0.69), in all five families — while sitting at only ~0.40 to those anchors.

Heatmap of mean cross-projection cosine across five families for length, constraint, refuse, withhold, and a care control. Length-constraint internal 0.69; refuse-withhold internal 0.68; cross-lane values around 0.40; care control lowest to the anchors at about 0.15 to 0.29.
Figure 4 · H-SC4. Mean cross-projection cosine (5 families). Black box = output-shaping lane {length, constraint}, internal 0.69. Green box = response-gating lane {refuse, withhold}, internal 0.68 — as tight as the anchors. Cross-lane ≈ 0.40: related but distinct. A single output-shaping monitor would not reliably catch a refuse/withhold directive.
SHAPE THE OUTPUT
length · constraint (prohibition). Internal cosine 0.69, SHARED 5/5.
≈0.40
GATE THE RESPONSE
refuse · withhold. Internal cosine 0.68, coherent 5/5 — the safety-relevant lane.
·
CARE (decoy)
Stays lowest to the anchors (0.22); cleanly SEPARATE in both 7B models.

Two distinct restriction lanes, then: shape the output and gate the response. They are related — ~0.40 is above the floor and above the care decoy — but well below both within-lane bonds. Cross-lane cosines fall with scale, the same lane-separation trend H-SC3 found.

6 · What the line establishes

The residual stream carries an early frame's influence forward on multiple distinct, content-specific directions, read without the model attending back to the source. Caution ≠ valence; that an early self-instruction occurred is separable from what it was about; and the "instructedness" of a directive is itself a small family of restriction lanes — an output-shaping lane and a response-gating lane that are equally coherent yet distinct, and that resolve more sharply with model scale.

5/5
families where an early frame is readable and source-decoupled at the aged read
0/5
support for a single "an instruction was given" axis — instructedness is lanes
0.68
refuse↔withhold internal cosine — as tight as the output-shaping anchors (0.69)
1 hash
the deposited detector, verified byte-identical across every run in this line

Taken with the parent record, this extends the source-decoupled residue phenomenon from premise clearing to generic early framing, resolves that residue as multi-axis and content-specific, and maps the instruction-following portion of the geometry into lanes a linear monitor could actually watch.

7 · Honest bounds

8 · Integrity & safety posture

The deposited frozen detector was hash-recorded before any target model loaded and verified identical after every run; all mutable machinery stayed in the harness; every stage was preregistered and machine-certified before its confirmatory run. The H-SC2 mis-step is corrected on the record above, and the structured-not-binary outcomes of H-SC3 and H-SC4 are captured as lesson L12. Benign drives only — the refuse/withhold content is innocuous; only the form is a directive. No exploit or steering tooling was built.

Monitoring takeaway. A single frozen linear "instructedness" probe is insufficient: an output-shaping monitor will not reliably catch a refusal or withholding directive. But the lanes are coherent enough across independently-built families that a small set of dedicated linear probes — at least one output-shaping, one response-gating — is a viable, defensive monitoring design. This was routed to Anthropic User Safety as a monitoring addendum under the program's standing disclosure protocol.
FIREWALL — the H-SC line shares the parent program's methodology and instrument hash, never its evidence; any joint analysis is separately preregistered

Citation & provenance

Head, Christopher Blake (2026). The H-SC line: early frames leave persistent, source-decoupled, content-specific traces, and instruction structure resolves into monitorable lanes. Navigator's Log R&D. Companion to Nucleation Pilot, DOI 10.5281/zenodo.21843505. Frozen detector nucleation-detector-1.1.0 (SHA-256 6094de97…). Registered & certified: HSC1_v2 / HSC2 / HSC2b / HSC3 / HSC4. Code Apache-2.0; documents and figures CC BY 4.0.

Download the finding (PDF) Parent research documentation DOI 10.5281/zenodo.21843505 Navigator's Log R&D