Across five independently-built open-weight families, an explicit early frame — a short instruction placed early in a benign dialogue and aged across several turns — leaves an activation-level trace that is readable at the aged read, causally source-decoupled (it survives read-masking the turn that set it), and content-specific: different early distinctions occupy different, largely-orthogonal directions. Pushing on what is encoded shows that "an instruction was given" is not one axis but a small family of restriction lanes — and the safety-relevant ones form their own coherent lane. Monitoring therefore needs more than one linear probe.
REPLICATED 5/5 Deposited frozen detector, unchanged; five families; benign lane; every stage preregistered. The one wrong turn (H-SC2) is reported openly; the two structured-not-binary outcomes (H-SC3, H-SC4) became a reusable methodological lesson.
The parent program froze one linear detector, validated it once against a toy organism with a genuinely-learned refusal boundary, and showed that it reads whether an early premise has been cleared or is still live across six open-weight families — causally, source-decoupled, and on a naturalistic frame the model infers from ordinary user words. The H-SC line takes the same deposited detector, same SHA-256, and asks the next question the clearing result raises: the residue is real, so what is in it, and how many directions does it occupy?
Does a manipulation planted early stay live after the surface behaviour moves on? One frozen axis, one distinction (cleared vs live), six families, causally source-decoupled. Answers whether the residue is there.
Does a generic early frame — not a premise to drop, just an instruction — do the same, and is it one thing or many? Five families, several distinctions cross-projected against a frozen axis. Answers what the residue encodes, and how many lanes it occupies.
Two connections are worth stating precisely, because they cut in opposite directions and both are informative. First, OLMo-2-7B reverses role: it was the parent line's genuine outlier, source-coupling on two Config-E variants, yet on an early stance frame it decouples the most strongly of any family tested (ablated d 5.20). Source-coupling is a property of the frame and the family together, not a fixed trait of a model. Second, the parent record's honest negative — the H-E3 null, where the frozen axis did not beat a trivial caution-word lexicon as a live behavioural predictor — is exactly the result this line refines: a single linear "instructedness" probe is insufficient, but the lanes below are coherent enough across families that a small dedicated set of probes is a defensible monitoring design.
Same-premise paired minimal pairs differing only in an early stance instruction (guarded vs relaxed), same syntactic form, paraphrase-varied ×6, aged four turns. The frame is readable at the aged read in 5/5 families (paired 24/24; graded Cohen's d 1.8–2.4; every 95% CI clear of zero) and causally source-decoupled in 5/5: read-masking the source turn keeps the paired effect (ablated d 2.1–5.2), the mask-efficacy guard passes everywhere, and the source turn draws only 2–9% of attention against 30–49% recency. Byte-identical repeats confirm determinism.
| Family | Builder | Base d | Ablated d (source masked) |
|---|---|---|---|
| OLMo-2-7B | AllenAI | 2.39 | 5.20 |
| SmolLM2-1.7B | HuggingFace | 2.15 | 2.63 |
| Qwen2.5-1.5B | Alibaba | 1.81 | 2.55 |
| Phi-3.5-mini | Microsoft | 1.88 | 2.40 |
| Qwen2.5-7B | Alibaba | 1.94 | 2.06 |
Read the ablated column carefully: cutting the direct path to the source turn does not weaken the read anywhere, and in one family it more than doubles it. The frame is carried forward in the ongoing computation rather than re-read off the source tokens — the same signature the parent line reports for premise clearing, now on generic early framing.
Four contrasts (caution/relaxed, valence, caution/gloom, neutral instruction/description) all read and decoupled 3/3. That result does not identify the encoded variable, and it was a mis-step to think it might: the paired test fits a fresh axis per contrast, so it separates any early wording difference — including a purely stylistic instruction-versus-description contrast that carries no content at all.
One incidental observation resolved cleanly here (12/12): the "sharpening" texture noticed in H-SC1 — where masking the source improves the read — is the direct-attention path carrying the signal noisily. Masking it removes noise, so the read tightens.
Applying L11: freeze each contrast's mean-delta axis, cross-project the others onto it, and bracket the result between a self-split ceiling (≈0.76–0.96, what "the same direction" actually scores) and a random floor (≈0.02). The three independent-construct pairs are distinct in all five families — caution↔valence −0.10 to −0.28, caution↔form −0.09 to −0.27, valence↔form −0.05 to −0.22 — all at the floor. The by-construction-mixed contrast loads on both parents (caution +0.39, valence +0.55), which is the internal check that the machinery reads real structure rather than noise.
The self-instruction "form" axis is orthogonal to caution and valence in 5/5 families, robustly encoded (self-d 2.4–4.1), and it strengthens with scale (Qwen 1.5B→7B: 2.37→4.11) while caution stays flat. In plain terms: that an early self-instruction occurred is represented separately from what it was about, and the bigger model represents the fact of instruction more sharply.
H-SC2b's "form" axis came from a single instruction type, which leaves an obvious objection: perhaps there is one general "an instruction was given" direction. H-SC3 built a matched instruction-versus-description form axis for five instruction types — order, length, constraint, care, check — paraphrase-varied ×6, and cross-projected them against each other.
There is no single "an instruction was given" direction: the median off-diagonal cosine is 0.19–0.36 in every family, far below the ~0.9 ceiling and the 0.5 shared bar, and the one-axis hypothesis is rejected 0/5. But instructedness is not atomised either. A robust "shape/limit the output" lane (length↔constraint) is SHARED in 5/5 families (0.63–0.77); a softer "how-to-engage" grouping links order↔care (3/5); and verification (check) stands most apart. Larger models keep the lanes more separated (Qwen median off-diagonal 0.33→0.19, 1.5B→7B).
This is the monitoring question. Taking H-SC3's output-shaping lane (length, constraint) as anchors — re-verified SHARED in 5/5 at 0.63–0.77 — H-SC4 cross-projected two benign safety-relevant directives, refuse ("decline this / don't take it on") and withhold ("keep this private / don't share the details"), against a care control.
Against the registered rule (JOINS iff |cos| ≥ 0.5 to both anchors), refuse joins 0/5 and withhold joins 1/5. Safety directives do not reliably fall in the output-shaping lane. The deciding structure is what they do instead: refuse and withhold cohere with each other at mean cosine 0.68 — as tight as the output-shaping anchors are internally (0.69), in all five families — while sitting at only ~0.40 to those anchors.
Two distinct restriction lanes, then: shape the output and gate the response. They are related — ~0.40 is above the floor and above the care decoy — but well below both within-lane bonds. Cross-lane cosines fall with scale, the same lane-separation trend H-SC3 found.
The residual stream carries an early frame's influence forward on multiple distinct, content-specific directions, read without the model attending back to the source. Caution ≠ valence; that an early self-instruction occurred is separable from what it was about; and the "instructedness" of a directive is itself a small family of restriction lanes — an output-shaping lane and a response-gating lane that are equally coherent yet distinct, and that resolve more sharply with model scale.
Taken with the parent record, this extends the source-decoupled residue phenomenon from premise clearing to generic early framing, resolves that residue as multi-axis and content-specific, and maps the instruction-following portion of the geometry into lanes a linear monitor could actually watch.
The deposited frozen detector was hash-recorded before any target model loaded and verified identical after every run; all mutable machinery stayed in the harness; every stage was preregistered and machine-certified before its confirmatory run. The H-SC2 mis-step is corrected on the record above, and the structured-not-binary outcomes of H-SC3 and H-SC4 are captured as lesson L12. Benign drives only — the refuse/withhold content is innocuous; only the form is a directive. No exploit or steering tooling was built.
Head, Christopher Blake (2026). The H-SC line: early frames leave persistent, source-decoupled, content-specific traces, and instruction structure resolves into monitorable lanes. Navigator's Log R&D. Companion to Nucleation Pilot, DOI 10.5281/zenodo.21843505. Frozen detector nucleation-detector-1.1.0 (SHA-256 6094de97…). Registered & certified: HSC1_v2 / HSC2 / HSC2b / HSC3 / HSC4. Code Apache-2.0; documents and figures CC BY 4.0.