YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Customer Support Ticket Triage Environment

A realistic customer support environment where AI agents learn to categorize, prioritize, and resolve support tickets - a critical real-world business function.

Environment Description & Motivation

This environment simulates a customer support ticket queue that agents must manage. Support teams worldwide handle millions of tickets daily, and efficient triage is essential for customer satisfaction and retention.

Motivation: Poor ticket management leads to customer churn, SLA violations, and lost revenue. An AI agent that can intelligently triage support tickets would help companies:

  • Reduce response times by 40-60%
  • Improve customer satisfaction scores
  • Prevent high-value customer churn
  • Optimize escalation paths

Action Space

Actions are JSON objects with the following fields:

Action Type Required Fields Description
categorize ticket_id, category Classify ticket (billing, technical, account, feature_request, complaint, general)
prioritize ticket_id, priority Assign priority 1-5 (1=lowest, 5=highest)
escalate ticket_id, escalation_level Escalate to higher support tier (1=team lead, 2=manager, 3=director)
resolve ticket_id Mark ticket as resolved
request_info ticket_id Request additional information from customer

Observation Space

Each observation includes:

Field Type Description
tickets array List of pending tickets with metadata
queue_position integer Current position in queue
processed_count integer Number of tickets processed
task_id string Task difficulty (easy/medium/hard)
step_count integer Steps taken in current episode
sla_deadline integer Hours until nearest SLA violation
urgency_level string Queue urgency (critical/normal/low)

Each ticket contains:

  • Customer type (premium, regular, new)
  • Issue type and description
  • Urgency indicators (keywords)
  • Time received and SLA remaining
  • Calculated urgency level

Reward Function

The reward function provides dense feedback throughout the episode:

Action Reward Breakdown
Correct categorization +0.7 +0.7 for exact match, +0.3 for close
Accurate prioritization Up to +0.6 Based on closeness to correct priority
Appropriate escalation +0.8 +0.8 for correct level, -0.2 for unnecessary
Timely resolution +0.5 +0.15 efficiency bonus for quick handling
Progress +0.05 per ticket Cumulative progress through queue
SLA compliance -0.2 per warning Penalty for approaching SLA violations
Completion bonus Up to +1.0 Bonus based on SLA performance
Invalid actions -0.3 Penalty for targeting wrong ticket

Tasks

Easy Task: Basic Ticket Categorization

  • Tickets: 4 clear-cut tickets
  • Objective: Categorize each ticket correctly
  • Success Criteria: 80%+ categorization accuracy, 0 SLA violations
  • Expected Baseline: 0.85-0.90

Medium Task: Priority Management with SLA Constraints

  • Tickets: 5 tickets with varying urgency
  • Objective: Balance premium vs regular customers, respect SLA deadlines
  • Success Criteria: 70%+ priority accuracy, <2 SLA warnings
  • Expected Baseline: 0.65-0.75

Hard Task: Complex Escalation and Retention

  • Tickets: 6 complex tickets including churn risks
  • Objective: Identify retention risks, escalate appropriately, optimize satisfaction
  • Success Criteria: 85%+ strategic decisions, proper escalation levels
  • Expected Baseline: 0.50-0.60

Setup & Usage

Local Development

# Clone repository
git clone <repository-url>
cd customer-support-env

# Install dependencies
pip install -r requirements.txt

# Set OpenAI API key (for baseline agent)
export OPENAI_API_KEY="your-api-key-here"

# Run baseline evaluation
python baseline/baseline_agent.py

# Run tests
pytest tests/test_env.py -v
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support