Runway Characters Is Live—but Web Sessions Stop at Two Minutes

Runway Characters remains available as a real-time conversational avatar product, but its simplest access route is deliberately limited. The current Characters documentation sets maximum sessions at two minutes in the web app, five minutes on the developer platform and 30 minutes through an API integration; it also lists a cost of two credits per six seconds, equivalent to $0.20 per minute.
This makes the product more concrete than it appeared around its initial release: there are now distinct testing and deployment routes, documented customization controls and published performance measurements. The central limitation remains unchanged, however—convincing real-time video is only one component of an avatar system, and a short browser conversation cannot establish the reliability of the language model, connected data or application actions around it.
Characters is a product within the GWM-1 family
The historical foundation predates the standalone Characters launch. Runway’s December 11, 2025 GWM-1 introduction presented an autoregressive model family built on Gen-4.5, generating video frame by frame, and divided it into separately post-trained variants for explorable environments, conversational avatars and robotics.
That separation is important. Characters is the audio-driven avatar branch of the family, not evidence that a single unified system simultaneously handles conversation, visual generation, environmental simulation and application logic. Its job is to turn incoming speech into an animated response with lip movement, facial expression, eye motion and gestures.
A complete interaction therefore extends beyond the video model. It may involve speech recognition, a language model, retrieval from supplied documents, text-to-speech, video generation, network transport and the host application. Calling the result a “real-time avatar” accurately describes the user-facing experience, but it does not identify which component produced an answer or caused a failure.
The browser demo and an API deployment answer different questions
The web app is suited to an initial visual check: whether a reference image produces a recognizable character, whether speech and facial motion remain aligned, and whether the delay is acceptable for a brief exchange. Its short session ceiling leaves little room to examine consistency over a lesson, support conversation or interactive presentation.
The developer platform adds custom settings and a longer test window. The API route is positioned for deployment and supports the longest documented sessions, while knowledge-base access and screen sharing are available outside the basic web route. Those distinctions make “30 minutes” a session-duration limit for an API integration—not a general promise of 30 free minutes for every account.
Session length also should not be confused with reliability. A longer call can expose visual drift, delayed responses and inconsistent behavior, but it does not by itself verify factual accuracy or successful backend actions. Those outcomes depend on the surrounding agent configuration as well as the avatar renderer.
One reference image lowers the creation threshold
A custom Character can begin with one reference image rather than a character-specific fine-tuning process. The supported range includes photorealistic people, animated figures, stylized characters and non-human subjects, which broadens the product beyond digital replicas of real presenters.
Appearance is only the first layer. A deployment can define a voice, personality and instructions, attach a knowledge base, expose approved tools and embed the resulting character in a web application. An external voice agent can also handle speech recognition, language-model responses and synthesized speech while Characters provides the visual output.
This modular design is commercially useful, but it complicates evaluation. An avatar can look fluid while delivering an incorrect answer from its language model or retrieval system. Conversely, a correct answer can arrive too slowly because of voice processing, network conditions or an integration outside the video generator.
The published latency figures need careful interpretation
The May 4, 2026 engineering account describes HD output at 24 frames per second from one reference image, an effective 37 milliseconds of model time per frame and 1.75 seconds of server-side turnaround from the end of a user’s speech to the first response frame; it also details autoregressive sequential streaming and concurrent operation of the diffusion transformer and video decoder.
These are measurements published by the developer, not an independent benchmark or a universal latency guarantee. Frame-generation time and conversational turnaround are different metrics: the first concerns video throughput, while the second includes the server-side voice-agent and video pipeline before network delay reaches the user.
The engineering description supports a stronger conclusion than the idea of a prerecorded animation loop attached to a chatbot. Frames are generated and streamed in sequence, while overlapping model stages helps the system meet its frame budget. Public material does not, however, justify attributing every visual reset or eye-motion artifact to a specific undocumented technique such as distillation, buffering or reused idle footage.
The consequential test is the whole agent, not the face
Runway’s entry into real-time avatars is no longer merely a model demonstration. Characters now has a defined product surface, multiple access routes and integration features that can connect an on-screen persona to documents and application functions.
For a serious assessment, four results must remain separate: visual stability, conversational delay, answer accuracy and successful tool execution. A brief web session can reveal obvious animation or timing problems, but it cannot establish how the system handles interruptions, unclear requests, unavailable data, failed actions or extended conversations.
The product’s most significant change is therefore not simply higher-quality avatar video. It is the packaging of single-image character generation as one deployable layer in a broader conversational system. That lowers the effort required to create an expressive on-screen identity, while leaving knowledge quality, permissions, safeguards and operational reliability to the system built around it.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.