The front door. These also appear under their use case below.
AI & ML interests
Agents make mistakes. Those mistakes are the most valuable training data there is. Only While. turns them into a dataset, post-trains an open model on it with SFT and RL, and proves the agent stopped repeating them on a held-out test.
Recent Activity
Text-to-SQL against a live warehouse. Graded by executing the query, not by reading it.
The same task at the same quality with fewer tool calls.
Training how an agent talks without changing what it knows. Judge scores the manner, code checks no required fact was dropped.
Constitutions and identity. A judge reads the reply against the principle, because code cannot grade manner.
Payment-intent data and the small models that verify what a shopper asked for before an agent acts on it.
The front door. These also appear under their use case below.
Training how an agent talks without changing what it knows. Judge scores the manner, code checks no required fact was dropped.
Text-to-SQL against a live warehouse. Graded by executing the query, not by reading it.
Constitutions and identity. A judge reads the reply against the principle, because code cannot grade manner.
The same task at the same quality with fewer tool calls.
Payment-intent data and the small models that verify what a shopper asked for before an agent acts on it.