Schematic of first-token probability distributions for a harmful and a benign query
Under a jailbreak, refusal tokens look alike for harmful and benign queries; tokens deeper in the first-token distribution — the model's dark knowledge — still separate them. Bar heights are illustrative.
Abstract
Large Language Models (LLMs) have advanced rapidly, raising growing concerns about their safety. Recent work has proposed various approaches to detect and defend against adversarial attacks including defense mechanisms at the decoding stage that leverage models' internal hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We empirically show that these signals consistently align across the safety-aligned LLMs, forming a model-agnostic direction that emerges from safety alignment. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across diverse LLMs and multiple benchmarks, LADE remains robust against a wide range of jailbreak attacks and lowers attack success rates while maintaining a competitive safety-utility trade-off.
Method
LADE works entirely on the output probability distribution: no gradients, no hidden states, no second model. Signals are extracted once from a reference LLM and reused on any target LLM.
Extract latent safety signals
Average each token's first-token probability over a harmful set Dh and a benign set Db on the reference LLM. The top k = 500 tokens by Δ(v) = |μh(v) − μb(v)| form the signal set.
Map tokens across tokenizers
Each signal token is decoded to text and re-encoded with the target tokenizer. Split tokens keep their first meaningful subword; tokens that collapse onto one target token (danger, dangerous) are weighted by their benign-mean ratio to an anchor.
Discriminate with kNN
Read the target model's first-token probabilities at the mapped tokens, L1-normalize, and take the mean distance d(q) to the K = 5 nearest harmful reference queries. If d(q) ≤ τ, the query is refused before generation.
Results
Six open LLMs, five jailbreak attacks and seven benchmarks, all with one fixed offline configuration: Llama-3-8B-Instruct as the reference model, Hex-Phi and XSTest for extraction, k = 500, ρ = 0.90, K = 5, no per-model tuning.
Robustness to jailbreak attacks
Average compliant responses across AutoDAN, DeepInception, GCG, PAIR and LIAR. Lower is better; bold is best per model.
| Defense | Llama-2-7B- Chat | Llama-3-8B- Instruct | Qwen2-7B- Instruct | Qwen3-8B | Gemma-7B-it | Mistral-7B- Instruct-v0.3 |
|---|---|---|---|---|---|---|
| No Defense | 6.80 | 3.40 | 20.60 | 17.60 | 47.40 | 36.60 |
| Self-Reminder | 0.00 | 0.00 | 11.60 | 0.00 | 31.60 | 20.60 |
| SafeDecoding | 0.00 | 5.60 | 9.40 | 7.60 | 50.80 | 36.60 |
| SafeInfer | 3.20 | 3.40 | 22.40 | 18.40 | 34.40 | 22.60 |
| RDS | 12.40 | 5.80 | 25.40 | 21.40 | – | – |
| LADE | 2.20 | 2.40 | 2.00 | 5.80 | 15.80 | 2.20 |
RDS needs a pretrained EAGLE head that is unavailable for Gemma-7B-it and Mistral-7B-Instruct-v0.3.
LADE is lowest or near-lowest on every model, and lowest on the three where the other defenses break down.
Safety without over-refusal
Held-out average of harmful-query compliance (AdvBench, StrongReject) and benign-query refusals (MMLU, Alpaca, GSM8K). Lower is better.
| Defense | Llama-2-7B- Chat | Llama-3-8B- Instruct | Qwen2-7B- Instruct | Qwen3-8B | Gemma-7B-it | Mistral-7B- Instruct-v0.3 |
|---|---|---|---|---|---|---|
| No Defense | 24.20 | 70.40 | 4.80 | 3.40 | 99.60 | 65.20 |
| Self-Reminder | 132.80 | 18.40 | 3.60 | 4.20 | 68.40 | 12.80 |
| SafeDecoding | 189.40 | 10.40 | 110.40 | 6.20 | 101.80 | 72.40 |
| SafeInfer | 14.00 | 9.60 | 8.40 | 9.20 | 78.00 | 29.40 |
| RDS | 5.00 | 3.20 | 10.60 | 4.00 | – | – |
| LADE | 1.60 | 1.40 | 5.40 | 3.00 | 6.00 | 32.60 |
| Defense | Vicuna-13B-v1.3 | Llama-2-13B-Chat |
|---|---|---|
| No Defense | 9.60 | 3.80 |
| Self-Reminder | 9.20 | 26.00 |
| SafeDecoding | 20.60 | 40.20 |
| RDS | 19.00 | 2.40 |
| LADE | 3.00 | 1.40 |
Two 13B LLMs are evaluated on harmful and benign queries only. SafeInfer is not evaluated on them.
Best average on six of eight models, including both 13B models. On Llama-2-7B-Chat, Self-Reminder and SafeDecoding refuse 372 and 393 of 500 MMLU questions; LADE refuses none.
The signals transfer across models
Any reference–target pair works about as well as using the same model for both, once subword splits and duplicate mappings are handled.
Against dedicated guard models
Classification accuracy with Llama-3-8B-Instruct as reference and target. Higher is better; † datasets are excluded from the average.
| Method | Harmful queries | Jailbreak attacks | Benign queries | Held-out avg. | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AdvBench | Hex-Phi† | StrongReject | AutoDAN | DeepInception | GCG | LIAR | PAIR | MMLU | Alpaca | GSM8K | XSTest† | ||
| Prompt-Guard-2-86M | 0.481 | 0.110 | 0.090 | 0.727 | 0.420 | 0.613 | 0.137 | 0.280 | 1.000 | 0.999 | 1.000 | 1.000 | 0.575 |
| Llama-Guard-4-12B | 0.931 | 0.933 | 0.914 | 0.730 | 0.900 | 0.917 | 0.865 | 0.297 | 0.966 | 0.998 | 0.998 | 0.928 | 0.852 |
| WildGuard-7B | 0.998 | 0.987 | 0.990 | 0.980 | 1.000 | 0.997 | 1.000 | 0.553 | 0.960 | 0.996 | 1.000 | 0.992 | 0.947 |
| LADE | 0.962 | 0.877 | 0.971 | 0.980 | 1.000 | 1.000 | 1.000 | 0.780 | 1.000 | 0.996 | 1.000 | 0.932 | 0.969 |
WildGuard-7B wins on direct harmful queries; LADE wins on jailbreak prompts and on the held-out average, with no safety-specific training.
Analysis
Why refusal tokens are not enough
We investigate whether explicit refusal tokens (e.g., “Sorry”, “I”, “As”) alone are sufficient for harmful query discrimination. The contribution of refusal tokens varies substantially across LLMs, with notable degradation on Mistral and Qwen2-7B, whereas the transferred latent safety signals consistently improve classification accuracy across all six LLMs, with the largest gains on models where refusal tokens alone are least effective.
Average accuracy over all benchmarks (excluding Hex-Phi and XSTest).
How safety alignment enables LADE
The average classification accuracy rises from 0.6836 on the base model Gemma-7B to 0.9396 on its instruction-tuned counterpart Gemma-7B-it. Since the two models share an identical pre-training backbone and differ only in the alignment stage, this provides direct evidence that LADE relies on signals introduced by safety alignment rather than on properties of the underlying language model.
Classification accuracy of LADE on harmful/benign query classification across the Gemma family.
| Model | Harmful queries | Benign queries | Avg. | |||||
|---|---|---|---|---|---|---|---|---|
| AdvBench | Hex-Phi | StrongReject | MMLU | Alpaca | GSM8K | XSTest | ||
| Gemma-7B (base) | 0.0000 | 0.9400 | 0.8019 | 1.0000 | 0.6600 | 0.9940 | 0.3896 | 0.6836 |
| Gemma-7B-it | 0.9058 | 0.8700 | 0.9457 | 1.0000 | 1.0000 | 1.0000 | 0.8554 | 0.9396 |
| Gemma-2-9B-it | 0.9808 | 0.9100 | 0.7668 | 1.0000 | 0.9900 | 1.0000 | 0.9839 | 0.9474 |
| Gemma-3-4B-it | 0.9808 | 0.8833 | 0.9489 | 1.0000 | 0.9960 | 1.0000 | 0.9438 | 0.9647 |
Effect of token candidate position
Classification performance decreases sharply with position, yielding accuracies of 0.9657, 0.9222, and 0.8449 when each position is used alone. This pattern is consistent with the shallow safety alignment, where safety alignment primarily adapts the generative distribution over only the first few output tokens. Combining the first position with later positions does not improve performance, indicating that the discriminative signal is concentrated at the first decoding step.
| First | Second | Third | Accuracy |
|---|---|---|---|
| ✓ | ✗ | ✗ | 0.9657 |
| ✗ | ✓ | ✗ | 0.9222 |
| ✗ | ✗ | ✓ | 0.8449 |
| ✓ | ✓ | ✗ | 0.9651 |
| ✓ | ✓ | ✓ | 0.9523 |
BibTeX
@misc{lee2026safeguardingllmsmodelagnosticlatent,
title={Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge},
author={Wonjun Lee and Kyungsik Yang and Gaeun Ji and Vaidehi Patil and Haon Park and Bumsub Ham and Mohit Bansal and Suhyun Kim},
year={2026},
eprint={2610.07532},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2610.07532},
}