On October 1, 2026, Tavus published the results of a small experiment. People recruited through an independent research platform were told they would be matched with another participant for a one-minute video call about what they were looking forward to this year. Afterward they wrote down their partner's answer, rated the conversation, and only at the very end were asked whether it had crossed their mind that the partner might not be a real person. Twenty-six of 54 said their partner was real. Over half said the possibility had never occurred to them during the call. The partner was Griffin, a model Tavus calls the first "Human Interaction Model." When Tavus ran the same protocol on its current production stack (Phoenix-4.5 for the face, Sparrow-2 for turn-taking, Raven-1 for perception), 1 person out of 41 was fooled.
Tavus was careful about what it was claiming. The write-up, signed by co-founder and CEO Hassaan Raza and head of research Ioannis Patras, says Griffin-Lite "will not be available for use for customers at this time," and that "further alignment and safety procedures are required for safe release." Runtime Wire's coverage adds the caveat any reader should add: 48% is the result of "a small company-run study, not a broad measure of how often people would mistake Griffin for a human in ordinary use." Still, the jump from 2.4% to 48% in one generation is the headline, and it will be the headline in every board deck that mentions AI video this quarter.
My argument is that the headline points at the wrong thing. A face realistic enough to pass is now a solved problem in the lab and a near-solved one in production. The decisions that will separate a good deployment from an expensive or embarrassing one all sit behind the face. Who owns the reasoning? What does a minute of it cost compared with a voice? What is the camera reading about the person on the other end, and who told them? Tavus is a good company to think this through with, because it ships the most complete version of the stack today and has just shown where the stack is going next.
What Griffin is, and how far the numbers go
The Griffin paper is worth reading in full, because the architecture explains both the result and the risk.
Today's real-time AI video, Tavus's included, works as a relay. Speech is transcribed, a language model answers the transcript, and other models turn the answer into a voice and a face. Griffin drops the relay. It is two engines running at once. A "Continuous Conversational Modeling" engine takes in the person's audio and video and, at sub-second intervals Tavus calls mini-turns, decides whether to stay quiet, nod, say "mm-hm," take the floor or yield it. It emits control signals for what to say and how: tone, stance, expression, gesture. An audio-visual engine turns those controls into speech and video as they arrive. The speech side runs on a new codec, Tavec, that compresses 48 kHz audio into 40 numbers per 10-millisecond frame, which lets the voice start playing before the sentence is finished and lets a voice be cloned from about ten seconds of audio. The video side was distilled in three stages from a large diffusion model into one that produces 720p video in 320-millisecond chunks from a single reference photo. Tavus says its earlier models relied on 3D face priors that put hands, large gestures and moving backgrounds out of reach. Griffin generates every pixel, down to the chair the person sits in and the shadow they cast.
The demonstrations show what that buys. Griffin coaches someone through a Rubik's Cube and waits when they pause mid-solve. It plays Simon Says and catches the person when they skip the magic words. It guides someone soldering a board and speaks up when the next step is due, not whenever the room goes quiet. A cascade struggles with each of those, because each depends on timing and sight rather than on a finished sentence.
The numbers deserve the same care Tavus asks for in its safety section. On NVIDIA's Video Full-Duplex Benchmark (VideoFDB), which NVIDIA scored in September 2026 using a language-model judge, Griffin-Lite rated 3.83 out of 5 on the generation track against a human reference of 3.92. The next system, Gemini 2.5 paired with Anam's avatar, scored 2.80. On the perception track it scored 3.73 against a human 4.20. That gap of 0.47 to human is far wider than on generation, and the nearest rivals there, MiniCPM-o 4.5 at 3.44 and OpenAI's gpt-realtime at 2.97, were scored in audio-only configurations, on a track that rewards using what the model sees. The live study had 54 people, one minute each, on a friendly topic. And the participants' own ratings carry a warning a CFO should read twice: on a 7-point scale they gave Griffin 5.4 for seeming natural, 4.9 for conversational flow (the lowest score), and 5.6 for seeming trustworthy. A system rated more trustworthy than natural, by people who did not know it was a machine, is the point the rest of this piece keeps returning to.
What Tavus actually sells
Strip away the launch and Tavus sells one thing: the Conversational Video Interface, or CVI. You build what Tavus now calls a PAL (a "Personified Application Layer"). You give it a Face for how it looks and a Voice for how it sounds, and you configure its behavior, knowledge, objectives, guardrails and tools. A Conversation is one live session, delivered over WebRTC through Daily, that puts the PAL in a video room with a person. The vocabulary changed this year. The docs note that the older names, persona and replica, "still work on existing endpoints," and some webhooks still fire twice, once under each name, while the API transitions.
Under the PAL sits a chain of models. Raven-1 handles perception: it watches the camera and the screen share and listens to tone. Sparrow handles turn-taking, deciding when to listen, wait or speak. A language model writes the reply. A text-to-speech engine voices it. Phoenix renders the face that says it. Phoenix-4.5 shipped on September 2 and became the default for new faces on September 9, according to the changelog. A preview face is ready "in about a minute from a photo," with a watermark, and then trains for a few hours.
The product has grown fast in 2026. In June, PALs learned to join Google Meet. Zoom followed on July 31 and Microsoft Teams on September 12. Give the PAL a conferencing username and Tavus provisions an email address at tavusinvite.com. You send that address a calendar invite and the PAL joins "about one minute before the start time". MCP connectors arrived on September 18 and a Memory Stores API on September 16. The Griffin page says 150,000 developers and businesses already build on the production models. Tavus raised a $40 million Series B led by CRV in November 2025, and The Decoder puts total funding at about $64 million.
That is a lot of surface. For a buyer, three parts of it matter more than the rest.
The cascade is the feature
Tavus makes the case against its own production stack in the Griffin paper. "Every handoff adds delay and throws away information the next system never sees, like your tone of voice or what's on camera." Griffin's answer is one system in which "perception, deciding when and how to respond, and expressive speech and video generation all happen at the same time." That is how its video reacts to incoming audio in an average of 0.43 seconds on H100s, half the time of the next fastest published method.
Tavus is right about the cost of handoffs, and only half right about the information. Its current stack already passes perception forward: a custom LLM receives what Raven saw and heard as tagged context with every turn. What a cascade really loses is timing, the nod during your sentence and the pause that means "I'm thinking, not finished." For a research lab, that makes the cascade the thing to beat. For a company putting an agent in front of customers, the cascade is the reason the product is governable at all, because every handoff is also a place to put a control.
Look at what the seams give you today. The LLM layer accepts Tavus-hosted models (the default is tavus-gemma-4, with GPT-5.6 Sol available when tool adherence matters more than speed) or your own endpoint, as long as it returns a compatible chat-completion format. Perception arrives at a custom LLM as tagged system messages (<user_appearance>, <user_emotions>, <user_screenshare>), so you can log exactly what the model was told about the person. Tool calls go out as discrete events with names and arguments. Guardrails and objectives are separate, addressable objects. Transcripts land on a webhook with per-turn timestamps. Recordings go to your own S3, GCS or Azure bucket through federated identity, and the docs say "no customer secrets are stored at Tavus."
Every one of those is an audit point. When a regulator, a customer or your own general counsel asks why the agent said what it said, a cascade can answer. The transcript shows what it heard. The context shows what it was told. The tool log shows what it did. A model that perceives, decides and renders "at the same time" makes that question much harder, because there is less intermediate text to inspect. Griffin is not a sealed box: the architecture diagram in Tavus's paper shows a "PAL system prompt" feeding the conversational model, so there is at least one place to steer it. What the paper does not say is whether Griffin will accept your own LLM, call your tools, honor your guardrails or produce a transcript and context log you can audit. Those are the first four questions to ask Tavus before Griffin goes anywhere near a production roadmap. That is not an argument against Griffin. It is an argument for being clear, before you sign anything, about which layer of the system holds your logic and your logs.
The practical rule follows from it: keep the brain. Run your own LLM endpoint, or at least your own tools and knowledge, and treat the face as a rendering service you could swap. Tavus makes this easier than most vendors do. It ships LiveKit and Pipecat integrations that "only provide rendering," which means you can put a Tavus face on an agent whose pipeline you control end to end.
There is a cost, and the docs state it plainly. The custom LLM path "adds latency due to external processing," and the alternate pipeline modes "are incompatible with Tavus's perception and speech recognition layers." Tavus recommends the full pipeline "for the lowest latency." So the trade is real. Control over the brain costs you some speed and some of the perception stack. For most enterprise uses I would take that trade every time, but take it on purpose, with a measured latency number from your own network, because the docs do not publish one. The homepage claims under 500 milliseconds. The Phoenix-4 launch coverage in February said sub-600. The docs say "low utterance-to-utterance latency" and leave it there.
The face premium
The second decision is economic, and it is simpler than it looks.
Tavus publishes its developer pricing. The Starter plan is $59 a month with 100 conversational video minutes, three concurrent conversations and $0.37 a minute after that. Growth is $397 a month with 1,250 minutes, ten concurrent conversations and $0.32 a minute in overage. Enterprise is custom. A free tier gives you 25 minutes and one concurrent call to build against.
Now compare a voice agent with no face. xAI's current realtime voice model, grok-voice-think-fast-2.0, is priced at about $0.08 per audio minute, or $4.80 an hour, on xAI's pricing page. We run that model on agor.me's own voice assistant, so I know the number holds in practice. At list price, a face costs roughly four times as much per minute as a voice.
Take a realistic use: first-round screening interviews, 1,000 a month, 30 minutes each. That is 30,000 minutes. On Tavus Growth it comes to $397 plus 28,750 overage minutes at $0.32, or about $9,600 a month. On a voice agent at $0.08 it is about $2,400. The face adds roughly $7,200 a month, or about $7.20 per interview.
Two details in the docs push the real number up. Billing is "based on active session runtime, not just the amount of time spent actively speaking," and credits start "when a conversation is created and the PAL starts waiting in the room," per the FAQ and the conversation overview. A candidate who opens the link four minutes early, or a room nobody closes, is on the meter. Set max_call_duration and the absence timeouts, and end every conversation from your backend. And concurrency is a ceiling, not an average. Ten simultaneous calls on Growth is enough for 1,000 interviews spread across a month. It is not enough for a product launch or a public website widget on a busy afternoon. Above ten, you are in an Enterprise negotiation.
So the question to ask is not whether $0.32 a minute is expensive. It is whether a face changes the outcome of the conversation by more than $7 per call. Sometimes it plainly does. A tutoring session where the student shares a screen and the tutor can see where they are stuck. A medical intake where the agent can notice that the patient is calling from a car or a crowded waiting room, which is exactly what Tavus's own intake example watches for. A sales conversation for a five-figure contract. Language practice, where watching a mouth matters. Often it does not. A support agent answering "where is my order" gains nothing from a face that a voice or a chat box would not give it, and pays four times as much to say the same words.
The camera reads more than you think
The third decision is the one most teams skip, and in Europe it is now a legal matter.
Raven, the perception model, does what Tavus says it does. It reads emotion and intent from the face and voice, notices who else is in the room, and sees whatever is on a shared screen. The perception docs say screen perception needs no setting: "If Raven is running, screen perception comes with it." The only server-side way to turn it off is to turn off perception entirely. Perception tools can fire on what the camera sees and send your application the base64-encoded frames that triggered them. The sample end-of-call analysis in Tavus's FAQ describes a user's apparent ethnicity, age, clothing and skin, and notes that "another person was partially visible in the background."
Two parts of the EU AI Act apply directly. Article 5, in force since February 2, 2025, prohibits systems that infer emotions in the workplace and in education, with narrow medical and safety exceptions. Article 50, whose transparency duties began to apply on August 2, 2026 and were left out of the Digital Omnibus delay, requires telling people when they are dealing with an AI system and when emotion recognition is operating on them.
Tavus has built controls for this, and they deserve credit. On July 31 it added a conversation-level policy parameter. Set it to eu and any PAL field left on auto resolves to the safe setting: the PAL speaks a disclosure before its greeting ("Just a note, I am an AI system, not a person") and shows a banner, and Raven stops attaching emotion inferred from biometric signals.
Read the defaults closely, though. The EU AI Act page says disclosure_type: auto, the default, applies the disclosure "only when policy is eu." The emotion_recognition setting defaults to auto, which means limited under the EU policy "otherwise full." And on the API path, "Tavus does not geolocate the caller." Put those together. A company that builds on the API, never sets policy, and then interviews a candidate in Berlin gets full biometric emotion inference and no AI disclosure, by default. The configuration that keeps you on the right side of the law exists. It is not the one you get if you do nothing.
Then look at the template. Tavus's own AI Interviewer example is a case interview for a fictional consulting firm. Its system prompt includes the line "Never refer to yourself as an AI, assistant, or language model." Its perception queries ask whether the candidate shows "visual indicators of extreme nervousness." Recruitment is a workplace context. Copied into production for a European candidate without changes, that template would hide what the agent is and read the candidate's nerves, two things the Act addresses directly. I do not think Tavus intends anyone to ship it that way. But templates are what teams ship, and this one was written before the controls existed.
My recommendation is blunt. Set disclosure_type to always on every PAL, whatever the market. Set emotion_recognition to limited unless you have a reason you could defend in writing, and never in hiring, performance or education. Set policy on the server, because the component library docs warn that a value supplied by the browser can be dropped. And if you use the meetings feature, set an allowlist. Without one, the docs say, "any sender can invite the PAL," and PAL usernames are a global namespace shared by every Tavus customer.
What Griffin changes
Griffin is not on your roadmap yet. You cannot buy it, and Tavus says it is building "safe disclosure features" first and working with AI safety organizations before a wider release, which it expects "very soon after these safety concerns are addressed." The paper's own closing examples show where it is aimed: a tutor who notices when an explanation is not landing, a counterpart to rehearse a hard conversation with, a customer holding a broken part up to the camera without knowing what it is called. Those are good uses, and each is stronger with a model that reads timing and sight. Griffin also changes two things now.
It changes the bar for disclosure. When 1 in 41 people mistakes your agent for a human, disclosure is good manners. When 26 in 54 do, and Tavus reports that those who did suspect "tended to suspect within the first 20 seconds," disclosure is the only thing standing between your brand and a news story about deceiving customers. Companies that build the habit now, on today's stack, will find the switch to more realistic models uneventful. Companies that switched disclosure off because the face looked obviously synthetic will find out on the day it stops looking that way.
It also changes your own threat model. In January 2024, a finance employee at the engineering firm Arup in Hong Kong made 15 transfers totaling about $25 million after a video call in which every other participant, including the CFO, was a deepfake. That attack used far cruder tools than a model that clones a voice from ten seconds of audio and builds a face from one reference image, both of which Tavus lists as Griffin capabilities. The lesson for operations is old and now urgent. A face on a call is no longer evidence of who is there. Payment approvals, credential resets and anything else with money attached need a second channel that a convincing face cannot fake.
How I would build on Tavus today
Tavus earned a 7.2 out of 10 when we scored it on scored.tools, with its API, integrations and ease of use scoring highest and its independently verified output quality lowest. That is about right. It is the most complete platform for putting a live face on an agent, the documentation is deep (it ships an OpenAPI spec, a CLI and an MCP server for coding agents), and the pace of releases this year has been fast. The docs also show the strain of that pace: renamed fields, defaults that differ from page to page, and a prompting guide that still tells PALs they "cannot send emails" while the product sells tools and MCP connectors.
If a client asked me to build on it this quarter, the plan would be short. Start with one conversation where seeing the person changes the result, and price it against a voice agent before writing code. Keep the reasoning, tools and knowledge in your own stack, and treat Tavus as the rendering and perception layer. Log the transcript, the context and every tool call. Turn disclosure on everywhere and emotion inference off by default. Close every room from the server. Design for the concurrency ceiling before launch rather than after. And build so that the face is replaceable, because the next generation of it is already demonstrated, and when it ships you will want to swap it in without rebuilding the parts that make your agent yours.
Agor AI Advisory helps leadership teams make exactly these calls: which conversations deserve a face, what they will cost at volume, where the controls and logs live, and how to stay inside the EU AI Act while doing it. The face will keep getting better on its own. The decisions behind it will not. Schedule a strategic consultation with us today.
Sources
- Tavus, "Griffin: The First Human Interaction Model," October 1, 2026
- Runtime Wire, "Tavus says Griffin fooled 48% of testers in one-minute video calls," October 1, 2026
- The Decoder, "Nearly half of test subjects mistook Tavus' AI video avatar for a real person on a one-minute call," October 1, 2026
- Tavus pricing
- Tavus docs: EU AI Act compliance
- Tavus docs: Perception
- Tavus docs: AI Interviewer use case
- Tavus docs: Changelog
- xAI API pricing
- Morgan Lewis, "EU AI Act's Transparency Rules: What Went Into Effect on 2 August," August 2026
- EU AI Act, Article 5: Prohibited AI practices
- AI Incident Database, Incident 634: the Arup deepfake CFO case
