Transluce Mental Health Suite Collection Data accompanying Transluce's Mental Health Behavior Report • 2 items • Updated 4 days ago • 3
CHIVE: Counterfactual Hypothesis Investigation Via Edits Collection Datasets and LoRA checkpoints for the CHIVE paper: Would this change your answer? Predicting the Effects of Prompt Edits on LLM Behavior in the Wild. • 13 items • Updated 28 days ago • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Diverse Deception Probes Collection Linear probes trained on diverse deception data to detect dishonest completions across model families (OLMo, Qwen, Gemma). • 5 items • Updated Mar 18 • 1
Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads Paper • 2607.01002 • Published Jul 1 • 19
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models Paper • 2605.11887 • Published May 12 • 19
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth Paper • 2605.25052 • Published May 24 • 14
1930 Coder Collection Fine-tuning the Talkie 13B 1930 model on agentic trajectories • 4 items • Updated May 5 • 4
(Some) Emergent Misalignment from Reward Hacking in RL Collection Model checkpoints from the project "(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL" • 228 items • Updated Jul 1 • 7
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents Paper • 2602.08964 • Published Feb 9 • 1
Faithful Persona-based Conversational Dataset Generation with Large Language Models Paper • 2312.10007 • Published Dec 15, 2023 • 11
Language Models Change Facts Based on the Way You Talk Paper • 2507.14238 • Published Jul 17, 2025 • 1