What happened
Tavus has introduced Griffin, a new class of AI it calls a 'Human Interaction Model' (HIM). Unlike traditional cascaded systems that chain speech-to-text, language models, and text-to-speech, Griffin uses a unified video-to-video architecture that processes audio and visual input simultaneously. The model continuously evaluates conversational state at sub-second intervals, allowing it to perform non-verbal actions like nodding, interrupting, or adjusting facial expressions in real time. In a live study conducted by the company, 48% of participants reported believing they were speaking with a real human, compared to a 2% pass rate for the company's previous systems.
Griffin operates as a two-part system: a Continuous Conversational Modeling engine and an Audio-Visual Generation engine. The modeling engine ingests audio and video to decide when to speak, listen, or react, while the generation engine uses a fast autoregressive diffusion (VDiT) to produce speech and video concurrently.
The system utilizes a convolutional autoencoder called 'Tavec' to map audio into continuous latents, enabling the model to stream audio packets as small as 10ms. This architecture allows the model to begin speaking or reacting before a user has finished their sentence, a capability Tavus claims is essential for natural 'back-channeling' and interruptions.
The video generation component uses a few-step autoregressive diffusion generator that produces 720p video in 320ms chunks. This allows the model to control not just the face, but the entire scene, including body gestures, chair movement, and background shadows, based on a single reference image.
Why it matters
Griffin represents a shift toward 'invisible' computing where AI interfaces mimic the nuances of human social interaction, such as timing, gestures, and emotional responsiveness. By moving away from turn-based, audio-only processing, the model aims to reduce the cognitive load of managing a machine, potentially making AI interactions feel more natural. The ability to generate 720p video in real-time chunks while maintaining conversational flow is a significant technical milestone in reducing for interactive AI agents. However, the reliance on a single reference image to generate full-body movement and background shadows raises questions about the potential for deepfake-style misuse, though Tavus claims to have a safety approach in place.
The 48% pass rate in the company's internal 'video Turing test' suggests a significant improvement in the perceived realism of AI avatars. By integrating visual context—such as reading a user's gaze or observing their environment—the model can provide more context-aware assistance, such as coaching a user through a physical task like solving a Rubik's Cube.
The technical shift to continuous, sub-second decision-making addresses a major pain point in current AI assistants: the 'dead air' and mechanical delays that occur when systems wait for a user to stop speaking before processing a response. By treating conversation as a continuous flow rather than a series of discrete turns, Griffin attempts to mirror the 'dance' of human communication.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
crm_get_transaction(id='4092').In AI, what are a model's "parameters"?
What to watch next
Tavus has released 'Griffin-Lite' as a research preview for a select group of early testers. The company has not disclosed specific pricing, broader access timelines, or the full technical specifications of the safety mentioned. Observers should watch for how the model performs outside of controlled research environments, particularly regarding its ability to maintain consistency during long-form interactions and its susceptibility to adversarial prompting.
Access to Griffin-Lite is currently limited to a select group of early testers. Tavus has stated that a more powerful version of the model will follow, but has not provided a roadmap for public availability or enterprise deployment.
The company has promised to release further details on its approach to safety. Given the model's ability to generate highly realistic, real-time video from a single reference image, the effectiveness of these safety measures in preventing impersonation or unauthorized content generation will be a critical area of scrutiny.