Findings of EMNLP 2026

On Mitigation of Subliminal Learning in Large Language Models

We study how subliminal traits evolve during fine-tuning and evaluate liminal training: an annealed KL-regularized method that constrains early drift from the base model.

Atsushi Yanagisawa*, Brendan Gho*, Rajendran Ramesh Babu Manoj Narender, Kevin Zhu & Antonio Mari

* Equal contribution · Algoverse AI Research
Can a student learn the intended task without also acquiring an unrelated trait from its teacher?See the setup ↓
01 / The problem

The dataset looks clean.
The signal isn’t.

Subliminal learning transfers behavioral traits through outputs that appear semantically unrelated to those traits—even after explicit mentions are filtered out.

01

A teacher has a trait.

A preference, response style, or other behavioral disposition.

→
02

It generates ordinary data.

Number sequences or correct math reasoning, screened for direct references to the trait.

→
03

The student can acquire it.

Fine-tuning may transfer an unrequested behavior alongside the intended task signal.

02 / Liminal training

Constrain early.
Anneal to zero.

Liminal training adds a KL-divergence penalty toward the frozen base model to the cross-entropy objective. The weight is held at λ₀ for the first epoch, then decays linearly to zero.

Training objective

Limit early distributional drift.

In the main experiments, training runs for three epochs with λ₀ = 1.0. The KL term is computed token-wise over completion tokens.

L(θ; t) = LCE + λKL(t) · DKL(θ₀ ‖ θ)
Liminal schedule · 3 epochs
hold λ₀anneal to 0
03 / Main results

Less measured trait transfer,
similar average task gains.

In the aggregate chain-of-thought results across five models, liminal fine-tuning reduced mean animal-trait acquisition relative to preference fine-tuning while producing a similar average GSM8K gain. Outcomes vary by model and trait.

Example trajectories

Liminal training stays closer to the control.

We measure trait-related probabilities at fixed checkpoints throughout fine-tuning. In this Qwen2.5-1.5B-Instruct example, preference fine-tuning produces larger increases for several animal traits, while liminal fine-tuning remains closer to no-preference fine-tuning.

Animal-trait probability trajectories for no-preference, preference, and liminal fine-tuning on Qwen2.5-1.5B-Instruct
Figure 1. Chain-of-thought fine-tuning trajectories across six animal traits, averaged across three fixed seeds.
CoT · five-model average

Mean ΔP(animal)

3.32 → 0.85
preference FT (%) → liminal FT (%)

Average change from baseline, using mean trait probability across training checkpoints.

CoT · five-model average

Mean ΔGSM8K

23.0 → 22.2
preference FT (pp) → liminal FT (pp)

Aggregate task gains are close, although the difference varies across individual models.

04 / KL timing

When KL is applied matters.

The main result raises a natural question: is the benefit due only to regularization strength, or also to when the constraint is applied? Holding peak weight fixed shows that timing changes the task–trait trade-off.

Normalized GSM8K accuracy versus normalized mean animal-trait probability for constant KL strengths and schedule variants

Early-weighted schedules lie on or near the empirical Pareto frontier; late-weighted schedules fall below the fixed-KL curve in the configurations studied.

Liminal

Hold λ₀ for epoch one, then decay linearly to zero.

Early anchor

Apply λ₀ for epoch one, then switch it off.

Constant KL

Sweeping λ traces a smooth task–trait trade-off curve.

Positive anneal

Increase the penalty from zero to λ₀ over training.

End anchor

Apply λ₀ only during the final epoch.

05 / Additional findings

Beyond animal-preference mitigation.

The paper also compares alternative mitigations, tests a response-style trait, and examines instability under no-preference fine-tuning.

Alternative mitigations are less consistent.

Paraphrasing still produces trait acquisition in many configurations. Freezing more layers partially reduces acquisition while increasingly reducing task learning.

The response-style experiment shows the same timing pattern.

For one Qwen2.5-1.5B experiment trained on French math traces, mean P(French) was 74.6% with standard fine-tuning and 5.5% with liminal training; GSM8K accuracy was 59.1% and 51.8%, respectively.

Fine-tuning on control data can still shift trait probabilities.

Initial trait probability strongly correlates with trajectory variation under the control condition in both CoT and number-sequence experiments.

06 / Boundaries

Limitations

Our experiments establish a consistent empirical pattern, but leave important questions open.

Traits

Most experiments concern animal preferences, plus one French response-style study. Subliminal transfer of misalignment was not studied because it could not be reliably reproduced in the models tested.

Tasks

The evidence covers number sequences and GSM8K chain-of-thought data. Broader datasets and task settings remain future work.

Mechanism

The experiments do not provide a mechanistic explanation of subliminal learning or why liminal training suppresses trait acquisition.

07 / Paper

Read, reproduce,
build on it.

Atsushi Yanagisawa*
Brendan Gho*
Rajendran Ramesh Babu Manoj Narender
Kevin Zhu
Antonio Mari
* Equal contribution · Algoverse AI Research
@inproceedings{yanagisawa2026liminal,
  title     = {On Mitigation of Subliminal Learning in Large Language Models},
  author    = {Yanagisawa, Atsushi and Gho, Brendan and
               Rajendran Ramesh Babu, Manoj Narender and
               Zhu, Kevin and Mari, Antonio},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  publisher = {Association for Computational Linguistics},
  year      = {2026}
}