DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents Paper • 2605.04808 • Published May 6 • 20
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection Paper • 2512.07533 • Published Dec 8, 2025 • 4
SecCodePLT: A Unified Platform for Evaluating the Security of Code GenAI Paper • 2410.11096 • Published Oct 14, 2024 • 13
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation Paper • 2505.23885 • Published May 29, 2025 • 1
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents Paper • 2505.05849 • Published May 9, 2025
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails Paper • 2502.05163 • Published Feb 7, 2025 • 22
DepthMaster: Taming Diffusion Models for Monocular Depth Estimation Paper • 2501.02576 • Published Jan 5, 2025 • 15
GRAPE: Generalizing Robot Policy via Preference Alignment Paper • 2411.19309 • Published Nov 28, 2024 • 47
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models Paper • 2410.10139 • Published Oct 14, 2024 • 51
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases Paper • 2407.12784 • Published Jul 17, 2024 • 51
Safe Reinforcement Learning via Hierarchical Adaptive Chance-Constraint Safeguards Paper • 2310.03379 • Published Oct 5, 2023
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation? Paper • 2407.04842 • Published Jul 5, 2024 • 55
DVERGE: Diversifying Vulnerabilities for Enhanced Robust Generation of Ensembles Paper • 2009.14720 • Published Sep 30, 2020
OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution Detection Paper • 2306.09301 • Published Jun 15, 2023 • 1
Mixture Outlier Exposure: Towards Out-of-Distribution Detection in Fine-grained Environments Paper • 2106.03917 • Published Jun 7, 2021
Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models Paper • 2404.02936 • Published Apr 3, 2024 • 3
Unsolvable Problem Detection: Evaluating Trustworthiness of Vision Language Models Paper • 2403.20331 • Published Mar 29, 2024 • 16
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression Paper • 2403.15447 • Published Mar 18, 2024 • 15