Beyond Evaluations: Trinitite on Building Trustworthy AI Agents

Sep 15, 2026
00:59:55

Loading video...

Show Notes

In this episode of the Convo AI World Podcast, Hermes Frangoudis sits down with Dustin Allen, co-founder of Trinitite, to explore why evaluating an agent is only part of making it reliable in production. Dustin traces Trinitite's origins to a productivity app for accountants and finance teams, where a damaged spreadsheet exposed the need to replay actions, restore data, and understand what went wrong. He explains why his team now treats AI delivery, governance, and security as connected parts of the same system, and how Trinitite uses deterministic inference to make its own assessments more repeatable and failures easier to investigate. The conversation gets practical about voice AI: where to place controls in a cascading pipeline, how asynchronous monitoring can preserve conversational flow, and why sensitive tool calls may need checks before an action proceeds. Dustin also discusses multimodal systems, developer tools, and the challenge of observing agents running locally on a device. For teams building voice agents or enterprise AI, the episode offers a grounded discussion of continuous monitoring, remediation, and the limits of guardrails, with an emphasis on learning from failures and keeping people involved in deciding which risks matter most.

Key Topics Covered

  • From productivity software to AI governance
  • Beyond evals: monitoring, policy enforcement, security, and AI delivery
  • Deterministic inference and repeatable assessments
  • Empty tool responses, workflow drift, and agent spending
  • Governing voice AI with asynchronous monitoring and tool-call checks
  • Cascading and multimodal pipelines
  • Guardian agents and local AI observability
  • Continuous assurance and human oversight

Episode Chapters & Transcript

0:00:12

Welcome: Dustin Allen of Trinitite

0:00:35

From productivity software to AI governance

0:04:03

Why evals alone are not enough

0:06:17

Monitoring without disrupting the conversation

0:08:59

Deterministic inference and clearer signals

0:13:23

Who is adopting AI governance?

0:16:20

What actually fails in production?

0:19:36

Turning failures into measurable remediation

0:21:05

From reactive fixes to continuous assurance

0:24:13

Why voice AI is harder to govern

0:31:03

Where governance fits in a voice pipeline

0:33:43

Governing multimodal systems

0:36:39

Bringing guardian agents to developers

0:40:54

Observing agents at the network and device layers

0:44:08

Connecting the governance stack

0:47:09

Evaluating how AI communicates

0:49:35

Where AI trust goes next

0:51:58

The limits of guardrails and the role of people

0:57:02

Should observability be a standard feature?

0:59:46

The wildcard: Physical AI

0:59:55

Closing and where to learn more

Click on any chapter to view its transcript content • Download full transcript

Convo AI Newsletter

Subscribe to stay up to date on what's happening in conversational and voice AI.

Loading form...
✓ Conversational AI news✓ No spam, ever✓ Unsubscribe anytime

Tags

#trinitite#dustin allen#hermes frangoudis#ai governance#ai observability#ai agents#trustworthy ai#deterministic inference#continuous assurance#voice ai#multimodal ai#guardrails#ai security#developer tools#conversational ai