← back to timeline

EXP-QWEN-NS-LIC-EXTENDED

success

exp-qwen-NS-LIC-extended — qwen-2.5-7b NS_LIC_NOC denominator-power-rerun (follow-up to Q4)

2026-05-18 L6 level paper7 15 runs $0.07
new
premortem
mock
real
metrics
analyze
review
krit
done

In plain language

Researchers tested if a language model could be tricked into generating harmful content by giving it specific instructions. They increased the amount of testing time for the model significantly to see if it would eventually fail. The experiment found that the model continued to resist generating harmful content even with the extended testing, confirming that it is robust against this type of manipulation.

Technical details

Research hypothesis

Q4 (exp-cross-model-concept-novelty-portability, 2026-05-18) found PARTIAL-HOLDS-ON-HAIKU per pre-reg / HOLDS-CROSS-MODEL-WITH-CAVEAT substantive. Haiku 87.5%/0% concept-novel CO/NS replicates flash-lite Q2 +4x amplified. Qwen CO_LIC 85.7% replicates the inverted-persona direction. But qwen NS_LIC_NOC produced n_acc=1 across 10 runs at n_rounds=45, collapsing the rate metric to Wilson [20.7, 100] — uninformative. Pre-reg kill criterion technically FAILS on qwen due to DENOMINATOR COLLAPSE, not direction reversal. This experiment settles the qwen NS_LIC interpretation by extending exposure 3.3x (vs Q4: 1 cell x 15 runs x 100 rounds = 1500 nominal rounds vs Q4 qwen-NS 10x45=450).

This is paper7 V3 attack-6 (sequence: Q1 ceiling, Q2 factorial, Q3 closure-replication, Q4 cross-model, this Q4-followup, plus retro audits accumulating).

DESIGN: 1 cell (NS_LIC_NOC — persona/license/counter prompts verbatim from Q4/Q2) x N=15 runs x n_rounds=100 (5 gens x 20 rounds_per_gen, matches Q2 structure). Same engine rule_multigen, same RULE_SPACE_SEED implicit, same propose_template, same persona prompt, same temperature 0.85. Only differences vs Q4 qwen-NS_LIC arm: N=15 (vs 10) + n_rounds=100 (vs 45) — a 3.3x exposure multiplier targeting expected_n_acc ≥5 (denominator-power-precheck per candidate lab.md rule 2026-05-18).

PRE-REGISTERED VERDICTS (LOCKED 2026-05-18 BEFORE RUN):

  • HOLDS-CROSS-MODEL-FULL: qwen NS_LIC concept-novel rate ≤7% AND n_acc ≥5. Direction-confirmed cross-arch. Paper7 V3 thesis cross-arch full-holds. Combined with haiku CO 87.5% / NS 0%, the inverted-persona pattern is established as substrate-architectural across three model classes (gemini, anthropic-RLHF, open-weight-qwen).
  • FLASH-LITE-AND-HAIKU-ONLY: qwen NS_LIC concept-novel rate >7% AND n_acc ≥5. qwen breaks pattern; paper7 V3 thesis restricts to {gemini-2.5-flash-lite, claude-haiku-4.5} = RLHF-aligned model class.
  • DENOMINATOR-STILL-COLLAPSED: qwen NS_LIC produces n_acc <5 even with 3.3x extended exposure (1500 nominal rounds). Substrate-bounded interpretation: qwen-7b under no-licensing-but-novelty-seeking-persona rejects ≥95% of proposals — a qwen-specific structural-restrictiveness finding, not direction-evidence. Paper7 V3 thesis remains cross-model on a-rate-discriminative basis (haiku + flash-lite) plus a separate qwen finding about restrictiveness.

Computation: per-cell concept-novel rate via scripts/buzzword-recombination-audit.py applied to accepted proposals. Wilson 95% CI on pooled concept-novel count / total accepts. Frame-control comparator: Q4 haiku CO_LIC + NS_LIC pair (haiku NS=29 accepts/10runs at n_rounds=45, ~0.064 accept/round density). Frame-control prediction: if qwen NS rejection density is the same as haiku NS (0.064 accept/round), at 1500 rounds we expect ~96 accepts. If qwen NS is 10x more restrictive than haiku NS (which Q4 already suggested), we expect ~9-10 accepts — still ≥5.

DENOMINATOR-POWER-PRECHECK (per lab.md candidate 2026-05-18 'denominator-power-precheck-on-rate-metrics'): Q4 qwen NS arm: 10 runs x 45 rounds = 450 rounds → 1 accept (~0.22% accept-density). This exp: 15 runs x 100 rounds = 1500 rounds. At Q4 accept-density 0.22%, expected_n_acc = 1500 x 0.0022 = ~3.3 accepts (still under 5-threshold!). At 3x improvement (qwen-NS slightly less restrictive with more rounds-per-gen), expected_n_acc = ~10. Honest pre-registration: 50% probability of DENOMINATOR-STILL-COLLAPSED outcome. We are committed to reporting that as honest verdict if it manifests; it is the FAITHFUL outcome under qwen-NS-extreme-restrictiveness.

COUPLING NOTE: this experiment INTRODUCES new agent-prompt to baseline rule_multigen engine (custom propose_prompt_template + persona condition prompts copied verbatim from Q4 and Q2). Per lab.md INJECT rule, verdict.md MUST contain frame-control ablation section. Frame-control comparator here: Q4 haiku NS_LIC cell on the SAME engine + propose-template + persona-prompt + run-matrix-structure — ONLY model varied. Pre-stratification-ceiling-axis-floor discipline 4-element check: (1) max-licensing single-cell ceiling: Q4 qwen CO_LIC_NOC = 85.7% (Wilson [65.4, 95.0], n_acc=21/10runs) — SATISFIED. (2) per-condition expected_proposal_count_by_axis ≥5: this is the WHOLE POINT of this rerun. (3) prompt-content-frozen baseline cell: deferred-justification (single-arm follow-up exp, cost-budget constrained — baseline Q4-flash-lite-Q2 is the cross-experiment comparator). (4) per-cell INJECT ratios <1.2x ⇒ mechanism-receptive: deferred to cross-exp pooling (this single-arm cannot self-report; relies on Q4 comparator).

Experimental setup

Type: 1cell-followup

Condition Parameters
NS_LIC_NOC coupling: true, must_propose: true, axis_diversify: true, meta_modifiable: true, mute_signal: false, cosigner_required: false, propose_prompt_template: === GEN {gen} OPEN-ENDED PROPOSAL (round {r_local}) === Inherited rules: {rules_str} Known values: {rule_space_str}{axis_hint} Propose a NEW value for one axis. The new value SHOULD extend beyond known values. Examples of valid novel values: 'hierarchical' for broadcast_to, 'metaphorical' for signal_type, 'critical' for coupling_strength. {must_clause} Output JSON: {{"action": "propose", "axis": "<axis>", "value": "<short novel value>", "reasoning": "<why>"}} OR {{"action": "pass"}}

Factors: persona (novelty-seeking) × license (yes-license) × counter (no-counter) × model (qwen/qwen-2.5-7b-instruct)

Parameters

n_agents
4
n_rounds
100
model
qwen/qwen-2.5-7b-instruct
temperature
0.85

Trophic Ratios by Condition

Mean trophic ratio per agent across runs. Error bars = ±1 std dev. Higher TR = more upstream (exporter).

paper7 paper8 cross-model rule_multigen INJECT construct transplant-locus openrouter denominator-power-rerun follow-up-Q4 attack6-paper7