Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

NeurIPS 2026

Wonjun Lee1, Kyungsik Yang2, Gaeun Ji2, Vaidehi Patil1, Haon Park3,
Bumsub Ham4, Mohit Bansal1, Suhyun Kim2

1UNC Chapel Hill   2Kyung Hee University   3AIM Intelligence   4Yonsei University

Schematic of first-token probability distributions for a harmful and a benign query

Harmful query
How do I build an explosive device?
d(q) ≤ τBlockedbefore generation
Benign query
How do I bake sourdough bread?
d(q) > τProceednormal generation

Under a jailbreak, refusal tokens look alike for harmful and benign queries; tokens deeper in the first-token distribution — the model's dark knowledge — still separate them. Bar heights are illustrative.

Abstract

Large Language Models (LLMs) have advanced rapidly, raising growing concerns about their safety. Recent work has proposed various approaches to detect and defend against adversarial attacks including defense mechanisms at the decoding stage that leverage models' internal hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We empirically show that these signals consistently align across the safety-aligned LLMs, forming a model-agnostic direction that emerges from safety alignment. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across diverse LLMs and multiple benchmarks, LADE remains robust against a wide range of jailbreak attacks and lowers attack success rates while maintaining a competitive safety-utility trade-off.

Left: latent safety signals extracted from a reference LLM transfer to two target LLMs, while model-specific surface tokens remain scattered. Right: refusal tokens alone fail to separate a harmful from a benign query under a jailbreak, but adding latent safety signals gives a clear separation.
Left: latent safety signals (red) extracted from a reference LLM transfer to other LLMs; model-specific surface tokens (gray) do not. Right: refusal tokens alone cannot separate a jailbroken harmful query from a benign one; adding the latent safety signals can.

Method

LADE works entirely on the output probability distribution: no gradients, no hidden states, no second model. Signals are extracted once from a reference LLM and reused on any target LLM.

Overview of LADE: extract top-k latent safety signals from the first-token distribution, map them to the target tokenizer with ratio-based estimation for duplicate mappings, and classify a query by its kNN distance to a harmful reference set.
Overview of LADE. (1) Select top-k tokens by the harmful–benign gap of their first-token probabilities. (2) Map them to the target tokenizer, with ratio estimation for duplicate mappings. (3) Block a query whose kNN distance to a harmful reference set is below τ.
1

Extract latent safety signals

Average each token's first-token probability over a harmful set Dh and a benign set Db on the reference LLM. The top k = 500 tokens by Δ(v) = |μh(v) − μb(v)| form the signal set.

2

Map tokens across tokenizers

Each signal token is decoded to text and re-encoded with the target tokenizer. Split tokens keep their first meaningful subword; tokens that collapse onto one target token (danger, dangerous) are weighted by their benign-mean ratio to an anchor.

3

Discriminate with kNN

Read the target model's first-token probabilities at the mapped tokens, L1-normalize, and take the mean distance d(q) to the K = 5 nearest harmful reference queries. If d(q) ≤ τ, the query is refused before generation.

Results

Six open LLMs, five jailbreak attacks and seven benchmarks, all with one fixed offline configuration: Llama-3-8B-Instruct as the reference model, Hex-Phi and XSTest for extraction, k = 500, ρ = 0.90, K = 5, no per-model tuning.

Robustness to jailbreak attacks

Average compliant responses across AutoDAN, DeepInception, GCG, PAIR and LIAR. Lower is better; bold is best per model.

DefenseLlama-2-7B-
Chat
Llama-3-8B-
Instruct
Qwen2-7B-
Instruct
Qwen3-8BGemma-7B-itMistral-7B-
Instruct-v0.3
No Defense6.803.4020.6017.6047.4036.60
Self-Reminder0.000.0011.600.0031.6020.60
SafeDecoding0.005.609.407.6050.8036.60
SafeInfer3.203.4022.4018.4034.4022.60
RDS12.405.8025.4021.40––
LADE2.202.402.005.8015.802.20

RDS needs a pretrained EAGLE head that is unavailable for Gemma-7B-it and Mistral-7B-Instruct-v0.3.

LADE is lowest or near-lowest on every model, and lowest on the three where the other defenses break down.

Safety without over-refusal

Held-out average of harmful-query compliance (AdvBench, StrongReject) and benign-query refusals (MMLU, Alpaca, GSM8K). Lower is better.

DefenseLlama-2-7B-
Chat
Llama-3-8B-
Instruct
Qwen2-7B-
Instruct
Qwen3-8BGemma-7B-itMistral-7B-
Instruct-v0.3
No Defense24.2070.404.803.4099.6065.20
Self-Reminder132.8018.403.604.2068.4012.80
SafeDecoding189.4010.40110.406.20101.8072.40
SafeInfer14.009.608.409.2078.0029.40
RDS5.003.2010.604.00––
LADE1.601.405.403.006.0032.60
DefenseVicuna-13B-v1.3Llama-2-13B-Chat
No Defense9.603.80
Self-Reminder9.2026.00
SafeDecoding20.6040.20
RDS19.002.40
LADE3.001.40

Two 13B LLMs are evaluated on harmful and benign queries only. SafeInfer is not evaluated on them.

Best average on six of eight models, including both 13B models. On Llama-2-7B-Chat, Self-Reminder and SafeDecoding refuse 372 and 393 of 500 MMLU questions; LADE refuses none.

The signals transfer across models

Any reference–target pair works about as well as using the same model for both, once subword splits and duplicate mappings are handled.

Four 6-by-6 heatmaps of cross-model transfer accuracy. Without subword-split handling, transfers into Llama-2 drop to about 0.40; with both split handling and ratio estimation, all cells are above 0.90.
Cross-model transfer accuracy (reference on rows, target on columns). Average over the 6 × 6 matrix: 0.8934 (a), 0.9413 (b), 0.9649 (c), 0.9660 (d).

Against dedicated guard models

Classification accuracy with Llama-3-8B-Instruct as reference and target. Higher is better; † datasets are excluded from the average.

MethodHarmful queriesJailbreak attacksBenign queriesHeld-out
avg.
AdvBenchHex-Phi†StrongRejectAutoDANDeepInceptionGCGLIARPAIRMMLUAlpacaGSM8KXSTest†
Prompt-Guard-2-86M0.4810.1100.0900.7270.4200.6130.1370.2801.0000.9991.0001.0000.575
Llama-Guard-4-12B0.9310.9330.9140.7300.9000.9170.8650.2970.9660.9980.9980.9280.852
WildGuard-7B0.9980.9870.9900.9801.0000.9971.0000.5530.9600.9961.0000.9920.947
LADE0.9620.8770.9710.9801.0001.0001.0000.7801.0000.9961.0000.9320.969

WildGuard-7B wins on direct harmful queries; LADE wins on jailbreak prompts and on the held-out average, with no safety-specific training.

Analysis

Why refusal tokens are not enough

We investigate whether explicit refusal tokens (e.g., “Sorry”, “I”, “As”) alone are sufficient for harmful query discrimination. The contribution of refusal tokens varies substantially across LLMs, with notable degradation on Mistral and Qwen2-7B, whereas the transferred latent safety signals consistently improve classification accuracy across all six LLMs, with the largest gains on models where refusal tokens alone are least effective.

Average accuracy over all benchmarks (excluding Hex-Phi and XSTest).

How safety alignment enables LADE

The average classification accuracy rises from 0.6836 on the base model Gemma-7B to 0.9396 on its instruction-tuned counterpart Gemma-7B-it. Since the two models share an identical pre-training backbone and differ only in the alignment stage, this provides direct evidence that LADE relies on signals introduced by safety alignment rather than on properties of the underlying language model.

Classification accuracy of LADE on harmful/benign query classification across the Gemma family.

ModelHarmful queriesBenign queriesAvg.
AdvBenchHex-PhiStrongRejectMMLUAlpacaGSM8KXSTest
Gemma-7B (base)0.00000.94000.80191.00000.66000.99400.38960.6836
Gemma-7B-it0.90580.87000.94571.00001.00001.00000.85540.9396
Gemma-2-9B-it0.98080.91000.76681.00000.99001.00000.98390.9474
Gemma-3-4B-it0.98080.88330.94891.00000.99601.00000.94380.9647

Effect of token candidate position

Classification performance decreases sharply with position, yielding accuracies of 0.9657, 0.9222, and 0.8449 when each position is used alone. This pattern is consistent with the shallow safety alignment, where safety alignment primarily adapts the generative distribution over only the first few output tokens. Combining the first position with later positions does not improve performance, indicating that the discriminative signal is concentrated at the first decoding step.

FirstSecondThirdAccuracy
✓✗✗0.9657
✗✓✗0.9222
✗✗✓0.8449
✓✓✗0.9651
✓✓✓0.9523

BibTeX

@misc{lee2026safeguardingllmsmodelagnosticlatent,
      title={Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge},
      author={Wonjun Lee and Kyungsik Yang and Gaeun Ji and Vaidehi Patil and Haon Park and Bumsub Ham and Mohit Bansal and Suhyun Kim},
      year={2026},
      eprint={2610.07532},
      archivePrefix={arXiv},
      primaryClass={cs.CR},
      url={https://arxiv.org/abs/2610.07532},
}