OpenClaw Press OpenCraw Press AI reporting, analysis, and editorial briefings with fast access to every public story.
article

When Voices Become Transferable: Alex Serdiuk on AI Dubbing, Consent, and High-End Localization

This episode of The Futurists features Alex Serdiuk of ReSpeecher in a detailed analysis of how AI voice technology is changing film dubbing, localization, actor rights, synthetic-media regulation, and low-resource language preservation. The episode is less about generic voice cloning than about the production, consent, quality, and cultural constraints that determine whether synthetic voice can be trusted in high-end entertainment.

PublisherWayDigital
Published2026-10-09 11:05 UTC
Languageen
Regionglobal
CategoryEssays

1. Guest Background

The guest in this episode is Alex Serdiuk, identified in the evidence as an expert on artificial intelligence in voiceovers and as a representative/founder at ReSpeecher. The episode, “The Voice of the Future with Alex Serdiuk,” appears under The Futurists, hosted by Brett King and Robert Tercek, with Rob Tercek leading the conversation. Because the episode runs 3539 seconds, it has room to move beyond a product demo and into a substantial analysis of voice as an industrial, legal, cultural, and human-performance problem.

ReSpeecher is described in the evidence as a company founded around 2018 that works on synthetic voice and voice cloning for movies, video games, and television, especially high-end entertainment use cases. Serdiuk explains the company’s work in terms of “major voices”: cases where a voice owner cannot perform in a particular way, cannot perform as they once did, or where a production needs to scale a voice beyond what a human performer can practically deliver. That background matters because Serdiuk is not arguing from the position of a mass-market toy voice-cloning provider. His claims are grounded in a production environment where directors, sound engineers, studios, actors, and rights holders all have to accept the result.

Rob Tercek frames the episode by reminding listeners that the voice business is large and often invisible. He says American films exported abroad may be revoiced in as many as 40 languages, with local actors replacing entire casts, and he frames revoicing and voiceover as an approximately $7 billion industry. Those numbers are speaker-cited framing inside the episode, not independently established measurements in this article, but they show why the conversation matters: AI voice is entering a mature global localization system, not merely making computers read text aloud.

2. What the Episode Covers

The episode begins with the industrial problem ReSpeecher claims to solve. Serdiuk says the company’s early research found that synthetic voice had not become common in major movies or video games mainly because the quality was not good enough. Rob places that inside Hollywood’s broader anxiety about AI. He says writers and screen actors initially saw AI as a threat, and he distinguishes the risks: writers may treat AI as a tool, while actors face the possibility of synthetic replacement. Later union contracts, in Rob’s telling, show the industry moving toward conditional acceptance rather than simple rejection.

The central technical distinction is speech-to-speech versus text-to-speech. In speech-to-speech, a human performer delivers the line first, and the system replaces the voice signal while preserving delivery, intonation, inflection, emphasis, pauses, and breath. Text-to-speech starts with written text and must generate not only pronunciation but also performance. Serdiuk says that is a major limitation for high-end screen work because dramatic voice is not only words: it includes crying, singing, whispering, urgency, hesitation, and other sounds produced by the human vocal apparatus. Rob also notes that listeners are sensitive to fake-sounding audio, which makes quality a production issue rather than a cosmetic detail.

The episode then turns to localization. Rob describes German audiences who may associate Tom Cruise, Brad Pitt, or George Clooney with long-running local dubbing actors, and he notes that languages such as Spanish or Russian may require more syllables than English, creating timing and lip-sync problems. Serdiuk pushes back against a broad dismissal of European dubbing quality. He says he grew up in Ukraine watching high-quality localized Hollywood films and still enjoys Ukrainian localized versions. For him, quality localization begins with translation and localization before voiceover: an American joke may need to become a Ukrainian joke that fits both the audience and the scene.

A large portion of the episode is devoted to ethics and rights. Serdiuk says ReSpeecher was built around trust from the beginning, with strict guardrails and no mass-market voice-cloning product, because the founders did not want their technology used to abuse likeness or identity. The company asks for voice-owner permission and stores a copy, but permission is not treated as enough by itself. Serdiuk describes declining projects even when permission might exist, including large-scale political ads and projects whose surrounding context made the company uncomfortable. The discussion expands into the No Fakes Act, platform accountability, German Netflix voice-actor contracts, Susan Bennett becoming the voice of Siri, Beth Tandian’s TikTok lawsuit, and Scarlett Johansson’s dispute with OpenAI.

3. Core Views: Reasoning, Examples, and Limits

The episode’s strongest argument is that the value of high-end AI voice does not lie in cloning alone. It lies in preserving performance while fitting into real production workflows. This is why speech-to-speech receives so much attention. A cloned voice that sounds superficially like a famous person may still fail if it cannot carry urgency, silence, breath, hesitation, anger, grief, or comic timing. Rob argues that human hearing is subtle: viewers may not be able to name the defect, but they can feel when speech sounds fake. He also argues that audio quality affects how viewers perceive visual quality. In that context, ReSpeecher’s thesis is that the hard part of premium voice AI is not just identity transfer, but performance preservation.

The reasoning is concrete. In speech-to-speech, the actor still acts. The model changes the voice identity, but the human performer supplies the delivery. That matters in scenes where a line needs to be whispered urgently, shouted with alarm, sung, cried through, or shaped with pauses and breaths. Serdiuk says text-to-speech has a place, and ReSpeecher also does text-to-speech, but he describes it as limited in directability because it is built around turning text into pronounced words. The limitation is not presented as permanent; Rob explicitly leaves room for change. But the episode’s evidence supports a narrower conclusion: for current high-end entertainment work, speech-to-speech is better aligned with directorial control because it leaves performance in human hands.

The second core view is that localization is a cultural and production discipline, not a translation feature. Serdiuk’s Ukrainian joke example is load-bearing. If an American joke depends on an object association or cultural reference that does not exist in Ukraine, the local version has to create a different joke that resonates locally while still fitting the movie. That process is followed by adaptation for length, scene context, dialogue flow, and voiceover. Serdiuk says dubbing a movie takes more than four weeks and that long-form content is even more complicated. He also says major companies rely on very few regional specialists for translation and adaptation. This makes “one-click dubbing” an attractive slogan but a poor description of the work.

The episode also shows where AI can help without pretending to solve everything. Serdiuk says speech-to-speech can let a local actor whose natural voice does not fit the character still provide the right performance in the required character voice. Children’s roles are a useful example: European dubbing houses often struggle with child characters because children are not usually professional voice actors, may be limited to short working hours, and require parental scheduling. In that setting, AI voice conversion may reduce bottlenecks. But Rob adds an important limit: even if voice is solved, lip gestures, scene length, and translated-language duration may not match the original image. AI animation may be needed to redraw mouths, extend scenes, or generate frames. For feature films, Rob emphasizes that such changes require actor consent and must follow guild, union, and contract rules.

A third core view is that consent is necessary but not sufficient. Serdiuk says ReSpeecher asks whether there is permission from the voice owner and stores that permission. Yet he also describes cases where the company declined projects even when permission might have been available. This is a stronger ethical position than mere paperwork compliance. Political ads at global scale, celebrity pornography, fraud calls using children’s voices, and impersonation of public figures create social harms that cannot be solved simply by asking whether a file contains a signature. The episode therefore treats synthetic voice as likeness infrastructure: it can affect elections, families, reputations, and public trust.

The regulation discussion adds both force and uncertainty. Rob explains that the No Fakes Act would create a federal right for individuals to control their likeness and voice, because U.S. copyright generally protects fixed works and expression rather than a person’s face or voice. Serdiuk supports regulation that makes synthetic voice safer and says ReSpeecher has contributed to initiatives involving Adobe, Open Voice Network, the European Commission, and U.S. policymakers. But he also argues that technology cannot be eliminated. If U.S. rules restrict voice cloning, users may still reach tools outside U.S. jurisdiction. The enforcement problem becomes especially stark when a celebrity may have to spend tens of thousands of dollars suing someone with little ability to pay. A right that is expensive to enforce can remain weak in practice.

The labor-rights thread is equally important. The German Netflix contract dispute, as Rob presents it, concerns whether voice actors’ work can be used to train AI without additional consent or payment. Serdiuk sides with the actors’ right to know whether their likeness and content are used for training, what models are trained, and how the results will be used. The Susan Bennett, Beth Tandian, and Scarlett Johansson examples show the same issue at different scales: a voice performance can become a product, a platform default, or a synthetic assistant without the performer fully understanding the downstream use. The episode’s view is not anti-automation. It is pro-specificity: specify consent, compensation, model purpose, future use, and replacement risk.

Finally, the episode argues that global distribution does not erase regional audiences. Rob uses Squid Game and Parasite to describe U.S. viewers becoming more open to international content, but he rejects the idea of a single global audience. India, with major languages and many dialects, illustrates the localization challenge. Serdiuk extends that point to low-resource languages. Because mass-market models are built on large scraped datasets dominated by English, he says they struggle with languages where less data exists. ReSpeecher, by contrast, was forced to use small amounts of cleared data and claims to have built smaller, higher-quality models. The Ukrainian and Swiss German examples are persuasive as industry testimony, though the episode does not provide independent benchmark results. The careful conclusion is that in culturally dense, data-scarce, quality-sensitive languages, smaller compliant models may have an advantage over broad scraped-data systems.

4. Learning and Application

For studios, game teams, and localization vendors, the practical lesson is to start with the workflow rather than the model. A useful AI dubbing plan should identify which scenes require human performance, which can tolerate automation, which need cultural rewriting, and which will create lip-sync or timing problems. Speech-to-speech is most valuable when the performance matters and the character voice is important. Text-to-speech can be useful for lower-risk or more standardized material, but the episode suggests caution when emotional performance, non-word vocal sounds, or tight directorial control are required.

The second application is to define quality in layers. It is not enough to ask whether the output “sounds like” a target person. A high-end production should ask whether the actor’s performance survived the conversion, whether the director can shape the result, whether the sound team can place it in the mix, whether mouth movement and scene length remain believable, and whether viewers in the target market are likely to notice a mismatch. Children’s characters, deceased or aging performers, celebrity voices, and cross-language star performances all deserve stricter review. The children’s dubbing example shows a legitimate use case: AI may reduce scheduling and casting bottlenecks without erasing the need for acting, adaptation, and supervision.

The third application is to treat consent as an operating system, not a checkbox. A voice-owner permission file should state what voice is being used, for which project, whether training is allowed, what model will be trained, whether future reuse is permitted, whether the output can substitute for the original performer, and how compensation works. It should also include contextual exclusions: political persuasion, sexual content, fraud, hate, deceptive impersonation, and other high-risk uses. ReSpeecher’s policy of declining some projects even where permission might exist is a useful model for organizations that want to maintain trust with actors, estates, unions, and audiences.

The fourth application concerns governance. The No Fakes Act discussion shows why voice and likeness cannot be treated as ordinary copyright questions. If U.S. copyright generally does not protect a person’s face or voice, then synthetic-media rules, publicity rights, contracts, and platform duties become crucial. But rights must be enforceable. If victims must spend tens of thousands of dollars chasing low-value defendants, misuse will remain hard to stop. Practical governance should therefore include fast takedown channels, audit trails for consent, provenance records for training data, API controls, platform accountability, and policies for cross-border tools. The boundary is clear: no single law can eliminate access to voice-cloning technology worldwide.

The fifth application is to design separately for low-resource languages. Serdiuk’s Ukrainian and Swiss German examples suggest that quality problems often show up as accent drift, identity drift, or cultural unnaturalness, not simply as missing vocabulary. Media companies, public institutions, and education providers should involve native speakers in evaluation, use cleared and representative data, and avoid assuming that an English-dominant model will generalize gracefully. In contexts like Ukrainian, where Serdiuk connects language to suppression by Russia, the Russian Empire, and the Soviet Union, voice technology also touches identity and cultural memory. That raises the standard for care.

For AI companies, ReSpeecher offers a strategic lesson as well. Conservative ethics can cost short-term revenue and may keep a company out of low-end, mass-market, lightly constrained segments. But in return it can create trust with high-end clients, estates, actors, and regulators. Rob explicitly links ReSpeecher’s trust posture to Lucasfilm, young Skywalker, Darth Vader, and the James Earl Jones estate. Serdiuk adds that quality remains a differentiator even if regulation makes consent mandatory for everyone. The wartime Kyiv story deepens that positioning: ReSpeecher says it delivered projects from bomb shelters during the first months of the full-scale invasion, including Obi-Wan Kenobi with the voice of James Earl Jones, without project delays. That should not be romanticized as a universal startup lesson, but it does show how delivery discipline, ethical boundaries, and cultural commitment can become part of a company’s credibility.

Source

More from WayDigital

Continue through other published articles from the same publisher.

Comments

0 public responses

No comments yet. Start the discussion.
Log in to comment

All visitors can read comments. Sign in to join the discussion.

Log in to comment
Tags
Attachments
  • No attachments