Beyond Evaluations: Trinitite on Building Trustworthy AI Agents
Loading video...
Show Notes
In this episode of the Convo AI World Podcast, Hermes Frangoudis sits down with Dustin Allen, co-founder of Trinitite, to explore why evaluating an agent is only part of making it reliable in production. Dustin traces Trinitite's origins to a productivity app for accountants and finance teams, where a damaged spreadsheet exposed the need to replay actions, restore data, and understand what went wrong. He explains why his team now treats AI delivery, governance, and security as connected parts of the same system, and how Trinitite uses deterministic inference to make its own assessments more repeatable and failures easier to investigate. The conversation gets practical about voice AI: where to place controls in a cascading pipeline, how asynchronous monitoring can preserve conversational flow, and why sensitive tool calls may need checks before an action proceeds. Dustin also discusses multimodal systems, developer tools, and the challenge of observing agents running locally on a device. For teams building voice agents or enterprise AI, the episode offers a grounded discussion of continuous monitoring, remediation, and the limits of guardrails, with an emphasis on learning from failures and keeping people involved in deciding which risks matter most.
Key Topics Covered
- •From productivity software to AI governance
- •Beyond evals: monitoring, policy enforcement, security, and AI delivery
- •Deterministic inference and repeatable assessments
- •Empty tool responses, workflow drift, and agent spending
- •Governing voice AI with asynchronous monitoring and tool-call checks
- •Cascading and multimodal pipelines
- •Guardian agents and local AI observability
- •Continuous assurance and human oversight
Episode Chapters & Transcript
Welcome: Dustin Allen of Trinitite
From productivity software to AI governance
Why evals alone are not enough
Monitoring without disrupting the conversation
Deterministic inference and clearer signals
Who is adopting AI governance?
What actually fails in production?
Turning failures into measurable remediation
From reactive fixes to continuous assurance
Why voice AI is harder to govern
Where governance fits in a voice pipeline
Governing multimodal systems
Bringing guardian agents to developers
Observing agents at the network and device layers
Connecting the governance stack
Evaluating how AI communicates
Where AI trust goes next
The limits of guardrails and the role of people
Should observability be a standard feature?
The wildcard: Physical AI
Closing and where to learn more
Click on any chapter to view its transcript content • Download full transcript
Convo AI Newsletter
Subscribe to stay up to date on what's happening in conversational and voice AI.