sync. (Sync Labs): the lip-sync engine behind Wav2Lip, scored 5.8 on the U365 CI-First Review

Status: Active | Last tested: 2026-09-25 (sync. (Sync Labs) as documented at sync.so in September 2026) | Re-check: trigger-based (max 6 months)
Active: the tool is current and recommended.
sync. (Sync Labs) is the lip-sync engine with the closest thing the category has to a research pedigree: the company's team wrote Wav2Lip, the 2020 open-source model that made mouth-to-audio generation work on footage in the wild, and the same paper introduced LSE-D and LSE-C, the two metrics the field still uses. The commercial product carries that lineage forward as a family of hosted models behind one API, a studio interface, two post-production plugins, and an MCP server. Nothing in this review suggests the technology has stopped working or that a competitor has overtaken it on the capability itself.
Re-check: trigger-based, maximum six months. The triggers are a new lipsync or react model release; a published lip-sync accuracy measurement from any source; a revision of the terms of service or the privacy policy, both dated 2024; a change to which plans carry the output watermark; or a reconciliation of the dubbing language count.
The re-check triggers above are not routine. Two of them exist because the vendor's own surfaces disagree with each other in ways a reader needs to know about: the resolution a mid-tier model actually produces, and how many languages the dubbing flow supports. A third exists because the legal documents that govern what the vendor may do with uploaded footage of an identifiable person are more than two years old while the product has shipped voice cloning, bulk dubbing, image-to-video, and assistant mode since then. Read Section 7c before you upload anything you would not want used in the vendor's marketing.

In this Tool Review
Status and Re-check
Status: Active | Last tested: 2026-09-25 (sync. (Sync Labs) as documented at sync.so in September 2026) | Re-check: trigger-based (max 6 months)
Status and Last Tested
Status: Active | Last tested: 2026-09-25 | Version reviewed: sync. (Sync Labs), as documented at sync.so in September 2026.
Active: the tool is current and recommended.
What Active means here. Active means current and recommended for the work this review describes: a recorded performance the team owns, a named human approver for every version, and a written record of the model, the cost and the approver. It does not mean the engine has been measured. No independent lip-sync accuracy figure exists for any commercial model in this family, and a reader who needs a measured guarantee should treat that as not yet available.
Re-check triggers:
A new lipsync or react model release from this vendor.
A published lip-sync accuracy measurement for any model in this family, from the vendor or from a third party, with a stated method.
A revision of the terms of service or the privacy policy. Both are currently dated 2024 and both predate voice cloning, bulk dubbing, image-to-video and assistant mode.
A change to which plans carry the output watermark, since that is the disclosure mechanism the operator cannot supply.
A reconciliation of the dubbing language count, which currently reads 29, 92 or 95 or more depending on the surface.
For detailed explanations of the CI-First evaluation terms used in this review, including the Humics Protection Badge and the AI Imposture Risk levels, see the Glossary at the end of this post.
The sync. naming, stated before anything else
Four names point at this product and only one of them is the vendor's current name. Get this wrong and you will read the wrong documents, price the wrong plan, or land on a competitor's page while searching for the vendor's own flagship model.
Name | What it is |
sync. | The current brand, written lower case with a full stop. The website is sync.so. |
Sync Labs | The name used throughout the documentation, the terms of service, and the API guides. |
Synchronicity Labs, Inc. | The legal entity. The terms of service name it as the contracting party and give a San Francisco address. |
synclabs.so | A domain the terms of service still name as the site where the service is provided. The current site is sync.so. |
Two collisions are worth knowing because they cost time rather than money.
The first is the model name. sync-3 is the vendor's flagship lip-sync model. Sync 3 is also the name of a Ford vehicle infotainment system that has nothing to do with any of this, and it is the higher-ranking search result for the phrase. A reader looking for documentation on sync-3 will meet car software. Worse, sync-3 and lipsync-2-pro also appear as selectable models inside ElevenLabs' own video workspace, which lists "Sync 3" and "Sync Lip Sync 2 Pro" as options in its model library. ElevenLabs reviewed separately in this series carries a Humics-Risky badge on its dubbing and cloning surface, and Sync Labs supplies at least part of the engine underneath that surface. The two products are in a supplier relationship on one side and a competitive relationship on the other, and the naming does not help a reader keep them apart.
The second is the model family itself. The vendor sells five named models, and the names do not sort by generation. lipsync-2 and lipsync-2-pro are one generation. sync-3 is a different architecture, not a version of lipsync-2. react-1 is an expression model with a separate job. lipsync-1.9.0-beta is the legacy fast model and is still sold. The scope of this review is the full family as it is sold on 2026-09-25, with sync-3 as the default, and the family table in Section 8 separates what changed between them.
One more scope note. This review is about the lip-sync and visual dubbing product. Sync Labs also ships a text-to-speech API and a voice cloning surface that feed the dubbing flow. Those surfaces are described here only where they change the lip-sync decision. The voice cloning question is governance territory and it sits in Section 7c, because the consent burden and the licence over the footage are the same document in both cases.
Tool Snapshot
Field | Detail |
Tool | sync. (Sync Labs), the hosted lip-sync and visual dubbing product of Synchronicity Labs, Inc. |
Category | Video and Creative Tool; Infrastructure and DevOps (sold as an API to developers) |
Primary job | Take an existing video and a new audio track or text and re-generate the speaker's mouth so it matches the new audio |
Vendor | Synchronicity Labs, Inc., 1700 Montgomery St, San Francisco, CA. Trading as sync. and Sync Labs |
Founded | 2023. The vendor's about page gives 2023 and a W24 Y Combinator batch |
Backing | Y Combinator, Google Ventures, Rebel Fund, Lobster Capital, Creator Ventures, per the vendor's about page |
Model family | sync-3 (default), lipsync-2-pro, lipsync-2, react-1 (expressive, subscription only), lipsync-1.9.0-beta (legacy) |
Architectures | sync-3 processes the full shot at once with built-in obstruction detection; lipsync-2 and lipsync-2-pro use two-second independent chunks per the model documentation |
Face resolution | 512x512 for lipsync-2 and lipsync-2-pro; the documentation describes sync-3 as 4K native. The lipsync-2-pro product page claims 4K support. See Section 8 |
Output | Up to 4K at 60 frames per second claimed for sync-3 on the model page. ProRes, MOV, MXF and DNxHD metadata handling in the app per the changelog |
Surfaces | Web studio, API (POST /v2/generate), Python SDK (syncsdk), TypeScript SDK, Adobe Premiere Pro plugin, DaVinci Resolve plugin, ComfyUI node, MCP server, ChatGPT plugin |
Input formats | MP4 video, WAV or MP3 audio, by public URL, direct upload (20 MB cap), or asset ID in the media library. sync-3 also accepts an image |
Max generation length | 20 seconds free; 1 minute Hobbyist; 5 minutes Creator; 10 minutes Growth; 30 minutes Scale; custom on Enterprise |
Rate limits | 100 requests per minute on the generate endpoint, 600 per minute on status reads, 10 per minute on the authentication endpoints, per the published rate-limit guide |
Concurrency | 1 on Free and Hobbyist, 3 on Creator, 6 on Growth, 15 on Scale, custom on Enterprise |
Free tier | 3 lip-sync generations per month, 20 seconds maximum, of which at most 1 may be sync-3 at 15 seconds; 10 text-to-speech generations per month; no card required |
Paid plans | Hobbyist 5 dollars per month, Creator 19, Growth 49, Scale 249, plus custom Enterprise, per the vendor's billing documentation |
Usage billing | Per output frame, quoted at 25 fps as a per-second rate. lipsync-2 from 0.04 to 0.05 per second; lipsync-2-pro from 0.067 to 0.083; sync-3 from 0.107 to 0.133; react-1 from 0.133 to 0.167; lipsync-1.9.0-beta from 0.02 to 0.025 |
Invoice thresholds | Usage accumulates and is invoiced when it reaches 6 dollars on Hobbyist, 20 on Creator, 50 on Growth, 250 on Scale |
Watermark | Free and Hobbyist outputs carry one. Creator and above do not |
Provenance feature | A proprietary watermarking and verification path on the vendor homepage that lets a third party check whether a video was modified using the vendor's technology. See Section 7c |
Batch processing | Up to 500 generations in one batch operation on Scale and above |
Refunds | Hobbyist and Creator, usage of 1,500 frames or fewer, base subscription invoice only and not the usage invoices, 5 to 7 business days, per the billing documentation |
Status | Active |
The Problem
Changing what a person said on camera has always been expensive, and the cost lands in one of three places.
The first is the reshoot. A script changes after the shoot, a phrase needs to be corrected, a name is pronounced wrongly, a claim needs softening. The presenter has to return, the lighting has to match, the wardrobe has to match, and the second take rarely matches the first. For a course module recorded once and updated quarterly, this is a permanent tax.
The second is the dub. A video that works in one language has a market in five or ten, and the standard answer is subtitles or a new recording with a new voice or a new presenter. Both change the product. Subtitles ask the viewer to read while a face speaks a language they can see but cannot hear. A re-recorded dub replaces the person's voice, and for face-to-camera work the mouth is the tell: a viewer watches the lips move to one set of sounds and hears another, and the brain registers an overdub even when the viewer cannot name what is wrong.
The third is the edit. Word-level changes to recorded speech used to mean re-cutting around the sentence, or covering the gap with a cutaway, or living with the wrong word.
The common thread is that the raw material exists and the performance exists. What does not exist is a mouth that matches the new audio. Every downstream problem, the cost of the reshoot, the wrongness of the overdub, the awkwardness of the cutaway, comes from that single missing piece.
There is also a problem the buyer does not usually name until later. The technology that fixes the mouth is the same technology that puts words in a real person's mouth, and the difference between a consensual dub and a fabrication is not technical. A team that adopts a lip-sync engine is adopting a governance obligation at the same time, and most product pages in this category do not help with it.
The Outcome
When the engine works, the mouth matches the audio and the result stops reading as an overdub. The visual result is a video where the speaker appears to say the new words, in the original take, under the original lighting, at the original angle, with the original performance around it. The vendor's own framing for the flagship model is that the leap takes the result from good enough to nearly indistinguishable.
Concretely, three outcomes are available to a U365 team that adopts this tool.
A recorded module is updated without a studio. The presenter records once, the script changes, and the changed sentences are re-generated against the original footage. The instructor's delivery stays the instructor's, and the class does not have to re-record for a corrected example.
A course or a campaign reaches three or four more markets without new talent. The original performer's voice is cloned from the same footage, translated into the target language, and re-synced. The viewer gets a version that looks shot in that language. This is the outcome the vendor's own workflow article describes as a single pass, and the vendor prices the whole thing per second of output rather than per seat.
A pipeline gets a mouth-fixing primitive inside it. Developers call one endpoint, receive a webhook, and store the result. Batch operations cover up to 500 videos on the higher plans, and the batch endpoint is the feature that turns lip sync from an editing step into a production line.
The outcome has boundaries worth stating in the same breath. The output is a new video, not an edited one: the rest of the frame is preserved and the mouth is generated, which means every frame is worth checking rather than spot-checking. The length cap is per generation, not per project, so long-form work is assembled from segments. The per-second cost is per attempt, so iteration costs money. And none of it removes the need for consent, disclosure, or an editorial check on what the new words say. Every outcome above is available only on footage the team has the right to re-author.
Who Should Use sync. (Sync Labs)
Localization and post-production teams with rights to their footage. This is the core buyer. A team dubbing its own library into several languages, with a presenter whose voice and likeness it already controls, gets the most value per dollar. The work that used to need a re-record, a new presenter, or an ADR session becomes an API call, and the per-second price is knowable before the job runs because the vendor ships a cost-estimation endpoint.
Software teams embedding lip sync into a product. The API, both SDKs, the webhooks, the batch endpoint and the MCP server together make this a genuine developer surface rather than an interface with an API bolted on. A team building a language-learning feature, a personalized video message product, or an in-app dubbing path can call it and forget it.
Marketing and brand teams producing multilingual campaign variants. One recorded message becomes several language versions with the original presenter's face. The honest caveat is that brand voice governance does not come with it: the output is only as on-brand as the script and the audio you supply.
Editors already working in Premiere Pro, DaVinci Resolve, or ComfyUI. The plugins put the model inside the tool the editor already uses, and the changelog records the direction of travel, which is one architecture behind studio, plugin, and API.
Video and media teachers. The tool is a clear demonstration object for a media studies or AI literacy class, because the output is designed to be indistinguishable from an unaltered recording and the only reliable check is a provenance mechanism the vendor switches off on paid plans. That is a real teaching case in one artefact.
Who should not use it, at least not in the common case. A team that needs to change a person's face rather than the mouth is buying the wrong class of tool: this product changes the mouth and preserves the face. A team with no source footage has nothing to feed it, because sync. generates no video from scratch apart from a newer image-to-video path that starts from a still. A team that wants a general content engine, a scriptwriter, a captioning tool, or a publisher will need all of those elsewhere. And a team that cannot answer who consented to the footage should stop before the first upload rather than after, because the licence it grants in the terms is broad enough that the answer matters to more than the project.
One further exclusion is worth naming because the framework scores the common case. A solo creator on the free tier gets three 20-second generations a month, one of which may be sync-3 at 15 seconds, and every free output carries a watermark. That is enough to evaluate the engine, and it is not enough to run a workflow on. Any professional use starts at the 19 dollar Creator tier, where the watermark comes off. Treat the free tier as a test bench, not a plan.
U365 Institutes Alignment
Institute ratings for this tool, confirmed by the University 365 Department of Academics as the academic owner. The primary institute is the one whose method value is greatest
The table below states the competency that remains with the Fellow once the tool is removed, not what the tool performs.
Institute | Rating | Why, stated as the competency that remains | The limit that holds the row |
UIC (Digital Communication, Marketing) | High (primary) | The competency is the publication rule for a re-timed or re-voiced version: which language markets a recorded message should reach, whether synthetic delivery is appropriate for this message to this audience, what the audience is owed by way of disclosure, what the version metadata records, and who reviews each version before it reaches a cohort. The tool re-times a mouth; the rule about when that is right to publish, and how the audience is told, cannot be delegated | The tool supplies the mechanism and not the judgement. It builds no audience analysis, no brand voice and no campaign craft, and no published UIC (Digital Communication, Marketing) programme assesses synthetic-media governance or the disclosure rule, so the competency is assessed in coursework rather than against a credential |
UIT (Technology, AI, Data Science) | High | The competency is integration engineering against a published specification: designing an asynchronous job flow and choosing between webhooks and polling, handling the terminal statuses the vendor documents, sizing a retry budget against a published concurrency table, and costing the job from a cost-estimation endpoint before spending anything. The tool is the third-party system being integrated, and the engineering knowledge is what a Fellow can explain without it | The tool teaches no programming or machine-learning competency of its own. It is a hosted service with weights that are not published, no training path and no reproducible evaluation surface, so nothing here is assessable as modelling. The engineering competency is assessable coursework and its credential anchors are adjacent rather than direct, as the Tool to Skill to Credential table below states |
UIB (Business Management, Entrepreneurship) | Medium | Two competencies. Usage-based cost appraisal: the per-second charge falls on every attempt, on a workflow whose normal condition is iteration, so the appraisal is a costed decision rather than plan arithmetic. And supplier-contract literacy: the perpetual and sublicensable licence over uploaded footage, the consent record the contract requires and does not define, and the clause forbidding a competing product built on the API, read as terms that constrain what a venture may build rather than what the vendor may do | The tool teaches no management, finance or entrepreneurship content of its own. Neither competency has a published U365 assessment home, and the contract reading is legal literacy applied to a purchase rather than a business discipline, so the row is coursework. A High rating would rest on coursework with no credential outcome behind it |
UID (Digital Design, UX/UI) | Medium, coursework observation only | The competency is critical evaluation of a motion artefact at frame level: watching an output designed to defeat casual review, checking the seams where the mouth is generated and leaving the middle alone, judging whether the occlusion handling held where a microphone crosses the face, and deciding whether the result meets the brief. That is a design-craft judgement the Fellow makes and the tool cannot make for them | The tool produces no design artefact, evaluates no design against a brief, and specifies no design method. It builds no UID (Digital Design, UX/UI) disciplinary competency, so the row stays at Medium rather than High, and no credential chain is mapped |
The test applied, stated once. A tool that performs a task cannot be chained to a competency in performing that task. It can be chained to the competency in judging, evaluating or directing it. The test is what remains when the tool is removed.
No institute is rated as having no relevance for this tool. The tool sits on a recorded human performance, which is the raw material of two institutes directly, it ships a first-class developer surface, and it raises a supplier-contract question that a venture has to answer. The four ratings therefore spread across High, High, Medium and Medium rather than clustering at Low. Where relevance is present but uncredentialed, this review says so rather than leaving it silent: no published U365 credential assesses the principal competency of any of the four rows.
Institutes alignment table
Institute | Rating | Why, stated as the competency that remains | The limit that holds the row |
UIC (Digital Communication, Marketing) | High (primary) | The competency is the publication rule for a re-timed or re-voiced version: which language markets a recorded message should reach, whether synthetic delivery is appropriate for this message to this audience, what the audience is owed by way of disclosure, what the version metadata records, and who reviews each version before it reaches a cohort. The tool re-times a mouth; the rule about when that is right to publish, and how the audience is told, cannot be delegated | The tool supplies the mechanism and not the judgement. It builds no audience analysis, no brand voice and no campaign craft, and no published UIC (Digital Communication, Marketing) programme assesses synthetic-media governance or the disclosure rule, so the competency is assessed in coursework rather than against a credential |
UIT (Technology, AI, Data Science) | High | The competency is integration engineering against a published specification: designing an asynchronous job flow and choosing between webhooks and polling, handling the terminal statuses the vendor documents, sizing a retry budget against a published concurrency table, and costing the job from a cost-estimation endpoint before spending anything. The tool is the third-party system being integrated, and the engineering knowledge is what a Fellow can explain without it | The tool teaches no programming or machine-learning competency of its own. It is a hosted service with weights that are not published, no training path and no reproducible evaluation surface, so nothing here is assessable as modelling. The engineering competency is assessable coursework and its credential anchors are adjacent rather than direct, as the Tool to Skill to Credential table below states |
UIB (Business Management, Entrepreneurship) | Medium | Two competencies. Usage-based cost appraisal: the per-second charge falls on every attempt, on a workflow whose normal condition is iteration, so the appraisal is a costed decision rather than plan arithmetic. And supplier-contract literacy: the perpetual and sublicensable licence over uploaded footage, the consent record the contract requires and does not define, and the clause forbidding a competing product built on the API, read as terms that constrain what a venture may build rather than what the vendor may do | The tool teaches no management, finance or entrepreneurship content of its own. Neither competency has a published U365 assessment home, and the contract reading is legal literacy applied to a purchase rather than a business discipline, so the row is coursework. A High rating would rest on coursework with no credential outcome behind it |
UID (Digital Design, UX/UI) | Medium, coursework observation only | The competency is critical evaluation of a motion artefact at frame level: watching an output designed to defeat casual review, checking the seams where the mouth is generated and leaving the middle alone, judging whether the occlusion handling held where a microphone crosses the face, and deciding whether the result meets the brief. That is a design-craft judgement the Fellow makes and the tool cannot make for them | The tool produces no design artefact, evaluates no design against a brief, and specifies no design method. It builds no UID (Digital Design, UX/UI) disciplinary competency, so the row stays at Medium rather than High, and no credential chain is mapped |
Tool to Skill to Credential
Six competencies this tool exercises, each with the nearest published U365 programme and what that programme does not publish. An adjacent anchor is useful to a Fellow who wants the neighbouring skill, and it is not a credential claim. No published U365 credential assesses any of the six competencies this tool exercises, and that is the finding rather than a catalogue defect: the tool removes a production requirement rather than teaching production, and U365 credentials assess what a Fellow can do rather than what a tool can do for them.
Tool skill | U365 competency | Nearest published programme, and what it does not publish | Institute |
Segmenting footage against the per-generation length cap and reassembling it, with a human check at the seams | Production craft and delivery discipline on a recorded performance | Video Production Specialist (60 days) publishes a Video Dialogue Editing module and the editing-application modules, and Adobe Specialist (60 days) publishes editing craft. Neither publishes an outcome in synthetic performance, lip sync or voice, so no credential recognises this competency | |
Evaluating a generated mouth at frame level against a brief, including occlusion and obstruction handling | Design-craft judgement applied to a motion artefact | Motion Graphics and VFX Expert (60 days) publishes compositing, rotoscoping and three-dimensional work, and 2D Animation Expert (30 days) publishes storyboarding and character development. Neither publishes an evaluation outcome for a synthetic performance, so the evaluation competency rests on coursework | |
Building an asynchronous generation integration: webhooks against polling, terminal statuses, a retry budget sized against a published concurrency table, and a cost estimate taken before the job | Integration engineering and evaluation against a published specification | AI Developer Specialist (18 days) publishes large language model application building and responsible algorithm design, and Python Developer (30 days) publishes programming craft. The certificate layer publishes the closest surfaces: MCP Server from Zero to Deployed (2 days, 1 US academic credit, 1.5 ECTS) and AI Agents and Workflows Automation with n8n (2 days, same credit). None publishes asynchronous job integration against a third-party rate-limit and concurrency specification | |
Costing a per-attempt, per-second charge on a workflow whose normal condition is iteration, and deciding when the charge is acceptable | Usage-based cost appraisal and a costed decision | Financial Analysis Specialist (30 days) publishes financial statement analysis, financial modelling and forecasting, and Business Analysis Professional (60 days) publishes benefits realisation and business process modelling. Neither publishes a usage-based or per-attempt pricing outcome | |
Reading the licence a lip-sync vendor takes over uploaded footage of an identifiable person, the consent record it requires and does not define, and the clause forbidding a competing product on the API | Commercial and legal literacy applied to a technology purchase | Entrepreneur (25 days) publishes a Business Law module, and Financial Analysis Specialist publishes the appraisal side. Neither publishes a technology-purchase contract-literacy outcome | |
Writing and enforcing the publication rule for a re-voiced or re-timed version, including the disclosure the audience is owed and the named reviewer per version | Editorial and publication rule applied to synthetic delivery | Content Marketing Specialist (30 days) publishes content strategy and live video production, and Social Media Marketing Manager (30 days) publishes copywriting and content creation for social. Neither publishes a disclosure or synthetic-media governance outcome, so the rule is assessed in coursework rather than against a credential |
Access levels, stated plainly. University 365 has three academic access levels: DISCOVERY, INSIDER and SUPERHUMAN. Specialised diplomas and certificates carry Basic, Foundation and Expert levels, where DISCOVERY Fellows can enrol in Basic-level programmes only, INSIDER Fellows in Basic and Foundation programmes, and SUPERHUMAN Fellows in all of them. University degree programmes carry a single Expert level and are open to SUPERHUMAN Fellows only. No degree chain is asserted for any of the six rows above, because the nearest anchors are diplomas and certificates.
Two curriculum gaps, recorded rather than filled. U365 publishes no assessment home for synthetic-media governance and the disclosure rule, which is the boundary document for a synthetic performance, the authorisation of what may be generated, the disclosure decision, and the named reviewer who stands between a model and another person. And it publishes no critical-listening assessment aimed at an output designed to defeat casual review, which is the specific skill this class of tool demands from the person who signs it off. Neither gap is a current programme and neither is presented as one.
How sync. (Sync Labs) Works
The engine takes two inputs and returns one output. The inputs are a video containing a face and an audio track. The output is a new video in which the speaker's mouth has been regenerated to match the audio, while the rest of the frame is preserved. The vendor's own description of the mechanism is that the model does not stretch or warp the original mouth: it generates new mouth movement for that specific face, in that lighting, at that angle.
Three surfaces call the same models.
The Studio is the browser app. You drop in a video, choose how to give it a voice, and generate. The vendor's own summary of the interaction is that there is no timeline and nothing to learn. A newer assistant mode goes further: you describe what you want, and the assistant picks the model, detects the speakers, and runs the generation.
The API is the developer surface. One endpoint creates a generation, a second returns its status, and webhooks can replace polling. The vendor publishes an OpenAPI specification, a cost-estimation endpoint, a batch endpoint for up to 500 generations, and a concurrency and rate-limit guide. An official Python SDK and an official TypeScript SDK wrap the same calls.
The editing surfaces are plugins. There is an installation guide for Adobe Premiere Pro, a plugin for DaVinci Resolve Studio, a ComfyUI node for custom pipelines, an MCP server that exposes the API to coding assistants and chat clients, and a ChatGPT plugin. The changelog records that studio, plugin and public API now run on one foundation, so a capability shipped in one appears in the others.
The model family, and what actually changed between versions
Read this table before you pick a model, because the family does not sort by generation and the price gap between the cheapest and the most expensive option is more than six to one.
Model | Architecture and what changed | Face resolution as documented | Speed as documented | Rate at 25 fps |
sync-3 | New architecture. Processes the full shot at once with a global read of the person, built-in obstruction detection, automated reasoning, native silent lip opening | 4K native | Fast (4 of 5) | 0.107 to 0.133 |
lipsync-2-pro | lipsync-2 plus diffusion-based super resolution for beards, teeth and fine detail. Adds a reasoning option. Slower than lipsync-2 | 512x512 with enhanced detail preservation. The product page separately claims 4K | Moderate (3 of 5), 1.5 to 2 times slower than lipsync-2 | 0.067 to 0.083 |
lipsync-2 | The zero-shot style-preserving model that replaced the legacy line. Two-second independent chunks for inference | 512x512 | Fast (4 of 5), about 1x baseline | 0.04 to 0.05 |
react-1 | Expressive lip sync with controllable emotions and head movement. Subscription only | 512x512, per the introduction page | No speed rating published | 0.133 to 0.167 |
lipsync-1.9.0-beta | Legacy fast model with generic mouth movement. Still supported but not recommended for accuracy | 512x512 | Fastest (5 of 5) | 0.02 to 0.025 |
Three things follow from that table.
First, the change that matters most between lipsync-2 and sync-3 is not resolution, it is the chunk boundary. lipsync-2 and lipsync-2-pro infer in two-second independent chunks, and the vendor's own caveat list says those models need natural speaking motion in the input. If your footage contains still frames where the speaker is not moving or speaking, lip sync does not work during those portions, even when audio is present. sync-3 builds a global understanding across the whole shot and generates all frames at once, and the documentation states that it can open silent lips to match audio, though the result is generic rather than speaker-style matched. If your footage contains any pause, any cutaway to a listening face, or any segment where the speaker is silent on camera, that single difference decides which model you can use.
Second, the resolution claim on the middle tier is not internally consistent, and it is the kind of inconsistency a buyer should resolve before spending. The model comparison table in the documentation puts lipsync-2-pro at 512x512 with enhanced detail preservation and reserves native 4K for sync-3. A question and answer on the same documentation set answers the resolution question by saying lipsync-2 and lipsync-2-pro generate faces at 512x512 and that sync-3 generates at 4K native, so resolution loss is not an issue. The dedicated lipsync-2-pro product page leads with a claim that the model now supports 4K and states that it preserves details even in 4K. The two readings are not reconcilable as written: one says the model works at 512x512 and the other says it supports 4K. The likeliest reading is that the product page means the model accepts a 4K source and returns it at 4K with a 512x512 face region upscaled, while the documentation means the face is generated at 512x512. A team buying for a 4K master should confirm which is true for its footage before it commits, and should treat sync-3 as the only model the vendor unambiguously documents as generating at native 4K.
Third, the pricing model is per output frame, quoted as a per-second rate at 25 frames per second. That distinction matters: a 60 frames per second source is billed on the output frames it produces rather than on a fixed per-second price. The same model can therefore cost noticeably more on a high frame rate master than the headline per-second rate suggests, and the vendor's own billing documentation says so plainly by stating that rates are quoted with 25 fps as a reference and that actual billing is per output frame regardless of the input frame rate.

What a job costs, worked through
Take the shape of a real U365 job: a ten minute recorded module that needs a French version and a Spanish version.
Under the Creator plan at 19 dollars a month, the maximum generation length is 5 minutes, so the module is processed in at least two segments per language. At the lipsync-2 rate of 0.04 per second and 25 frames per second, ten minutes of output is 600 seconds of charge, or 24 dollars of usage per language. sync-3 at 0.133 per second is 79.80 dollars per language, and lipsync-2-pro at 0.083 lands at 49.80. Two languages on lipsync-2 costs about 48 dollars of usage on top of the subscription. The Scale plan applies a 20 per cent usage discount and raises the generation cap to 30 minutes, so the same job costs less per second but requires 249 dollars a month.
The vendor makes this easy to check before you commit, because the cost-estimation endpoint and the in-app estimate both exist. Use them. Every re-run is billed, and dubbing work always has re-runs.

One billing mechanic is worth understanding because it surprises people. Usage is not prepaid and there are no credits. Usage accumulates until it reaches the invoice threshold for your tier, at which point the card is charged automatically, so a heavy month can produce several invoices: the vendor's documentation gives the example of the Creator plan creating an invoice each time accumulated usage reaches 20 dollars. The thresholds are 6 dollars on Hobbyist, 20 on Creator, 50 on Growth and 250 on Scale. The documentation also records an unpaid-usage pause in which new generations stop temporarily between invoices, and says the pause usually clears within minutes as long as the payment method is valid.
Speed and latency, as claimed
No latency service level is published. The documentation's quickstart says most generations complete within a few minutes depending on model and input length, and a vendor article on translating YouTube videos says the transcribe, translate, voice-clone and re-sync pass usually completes in under three minutes. Model selection changes the wait: the documentation rates lipsync-2-pro as moderate speed and 1.5 to 2 times slower than lipsync-2, and it notes that enabling obstruction detection on the two chunked models improves face detection in complex scenes but comes with slower generation speeds. Nothing in the documentation promises a turnaround time for a given length of footage, and the product is explicitly not suitable for real-time or low-latency use.
Independent measurement of lip-sync accuracy
There is none that applies to these commercial models. The vendor's own research did the field its most useful service: the Wav2Lip paper introduced LSE-D, the lip-sync error distance, and LSE-C, the lip-sync error confidence, and released ReSyncED, a benchmark set of real synced videos, exactly because the benchmarks in use could not distinguish a well-synced model from a badly synced one on real footage. The vendor's blog is right about that, and the metrics are still used across the literature.
What does not exist is a published accuracy figure for sync-3, lipsync-2-pro, lipsync-2 or react-1 on any independent benchmark. No LSE-C or LSE-D number, no human preference study, no third-party head-to-head with a stated methodology. The vendor's model comparison table has an Accuracy row, a Speed row, an Identity Preservation row, a Teeth row, a Face Blending row, a Beard row and a Pose Robustness row, and every one of them is empty apart from the qualitative text in the best-for column. The one quantitative speed statement in the documentation is a relative benchmark, 1.5 to 2 times slower than lipsync-2 for lipsync-2-pro, alongside a five-point qualitative scale for the rest.
Two consequences follow, and Section 11 scores on them. Quality is capped by the absence of independent measurement rather than by measured weakness. And a buyer cannot select between the five models on evidence, only on the vendor's qualitative prose, which is what drives a per-second price difference of up to six to one.
Where the vendor's own surfaces disagree
Four spreads are live at the same time, and a reader will combine them unless they are separated. None of them is a scandal, and all four are the kind of thing a buyer should price correctly.
Thing | Surface A | Surface B | Surface C |
Rate limit on the generate endpoint | The current documentation introduction states 100 requests per minute | The dedicated rate-limit guide states 100 requests per minute | An older published introduction page states 60 requests per minute on the same endpoint |
Lip-sync languages | The sync-3 model documentation states 95 or more languages | The changelog states built-in dubbing now supports 92 languages | The earlier dubbing changelog states one-step dubbing in 29 languages, and a vendor article on translating YouTube videos states 95 or more |
lipsync-2-pro face resolution | The documentation model table states 512x512 with enhanced detail | The same documentation's question and answer states 512x512 for both lipsync-2 models | The lipsync-2-pro product page states the model now supports 4K |
Growth plan price | The vendor's billing documentation states 49 dollars a month | A third-party review states 99 dollars a month for the same tier | The vendor's pricing page carries the current figure |
The language spread has a readable explanation, and the explanation is useful. The 29-language figure comes from the first release of the built-in dubbing flow, which the changelog describes as powered by an external voice provider and which a vendor blog post confirms uses that provider's voice technology. The 92-language figure is the newer dubbing engine with automatic speaker detection and lossless audio. The 95 or more figure is the language coverage attributed to the model family and repeated in a vendor article. So the three numbers describe three different things: the first-generation dubbing flow, the current dubbing flow, and the model's own language reach. The documentation for the flagship model answers the language question with 95 or more languages and describes that as the same broad language coverage as previous models, which quietly contradicts the model release note describing sync-3 as a fundamentally new architecture. If you are choosing a model partly on language coverage, ask the vendor which number applies to the dubbing path you intend to use.
The rate-limit spread is the least consequential of the four and the easiest to check at runtime, because the current guide is the one the service follows. The resolution spread is the one that can cost money.
Getting Started with sync. (Sync Labs)
A working first session takes about fifteen minutes. The steps below follow the vendor's own quickstart and free-trial documentation.
Create an account at the vendor's sign-up page. No card is required for the free tier.
Confirm what the free tier gives you before you plan anything. Three lip-sync generations per month, 20 seconds maximum each, of which at most one may use sync-3 and then at 15 seconds maximum. Ten text-to-speech generations per month. All lipsync family models are available including lipsync-2-pro. Unused generations do not roll over. Free outputs carry a watermark.
Pick your surface. The Studio needs no code. The API needs an API key, created in the dashboard settings.
If you are using the API, install an SDK. The Python package is installed as syncsdk and requires Python 3.8 or later. The TypeScript package is installed as the vendor's scoped SDK package and requires Node.js 18 or later. Set the API key as an environment variable; forgetting that is the vendor's own first listed pitfall.
Pick a model deliberately, not by default. lipsync-2 is the general-purpose choice. lipsync-2-pro is the premium detail choice. sync-3 is the default and the only one documented as native 4K with automatic obstruction handling. lipsync-1.9.0-beta is the cheap fast legacy path. react-1 needs a paid plan.
Prepare inputs to the documented limits. MP4 video and WAV or MP3 audio. Inputs can be a public URL, a direct upload capped at 20 MB, or an asset ID from the media library. Audio above 300 seconds is rejected, per the vendor's own pitfall list.
Understand that generation is asynchronous. The create call returns an identifier and the job finishes later. Poll the status endpoint or register a webhook. The terminal statuses are completed, failed and rejected.
Estimate the cost first. Use the cost-estimation endpoint or the in-app estimate. Every attempt is billed.
Run one short generation on real footage you own. Face-to-camera clips work best for a first test, and the vendor's own guidance is to watch the mouth rather than the waveform.
Check the frame rate of your master and the resulting charge. The per-second rates in this review assume 25 frames per second, and billing follows the output frame count.
Two setup decisions belong in the plan rather than in the tool, and both have a cost if they are skipped.
Decide your disclosure rule before the first upload. The vendor's own article on safe AI dubbing says the difference between a dub and a deepfake is authorization, and advises keeping a record of permission whenever you dub a voice that is not yours. The terms of service put the whole consent burden on you and require no record. Write the rule down anyway, and keep the records, because the licence you grant in the terms is broad and perpetual. Section 7c quotes the operative text.
Decide whether you need the watermark off, and understand what that trades. Free and Hobbyist outputs carry the vendor's watermark. Creator and above do not. The vendor also offers a verification path on its site that lets a third party check whether a video was modified using its technology. Paying 19 dollars a month turns that signal off on your own output. That is the right trade for a disclosed dub and the wrong one for a message that is meant to be attributed to a real person without a disclosure.
Real Workflows
Three workflows follow, each with a verification checklist. The verification checklists are not optional decoration: they operationalize the Executive Safeguard, and they are where most of the risk in this class of tool actually sits.
Workflow 1: Localising a recorded U365 module into a second language
The job: a recorded module taught in English needs a French version for a francophone cohort, with the same instructor's face and voice.
Steps. Confirm in writing that the instructor consents to the re-voiced version and that U365 holds the rights to the footage. Export the master at the frame rate you intend to deliver. Segment the module at the Creator plan's 5-minute generation cap or higher tier. Clone the instructor's voice from the same footage, or use text-to-speech on the translated script, through the dubbing flow or a separate voice step. Run the lip sync per segment on sync-3 if the footage contains pauses, obstructions or close-ups, and on lipsync-2 if it is a clean continuous talking head. Attach the dubbed audio track in the hosting platform as a separate language track rather than replacing the original. Label the version as AI-dubbed in the description and, at U365's editorial standard, in the module metadata.
VERIFICATION CHECKLIST for Module localisation:
[ ] Multi-Model Check: run one segment on lipsync-2 and the same segment on sync-3, and compare the two outputs side by side on the same monitor. If the cheap model matches the expensive one on this footage, use the cheap model for the rest of the module.
[ ] External Source: have a native speaker of the target language watch the full version once and report any line where the mouth, the stress or the emotion reads wrong. A fluent viewer catches what no metric in this category publishes.
[ ] Human Review: the instructor, or a delegate who knows the instructor's delivery, approves the final version before it reaches a cohort.
[ ] CI-First Test: can the team explain why each segment used the model it used, and what the re-run budget was, without the tool's dashboard? If not, the decision was the tool's, not the team's.The trap in this workflow is the model choice. lipsync-2 is a quarter of the price of sync-3 and produces excellent results on clean, continuous, face-to-camera footage. It stops working on still frames, on cutaways to a listener, and on the moments where a microphone crosses the mouth. A module with slides intercut between the speaker has still frames, and the cheap model will visibly fail on exactly those moments. Run the comparison on one segment before you commit a whole module to either model.
Workflow 2: Correcting recorded content without a reshoot
The job: a recorded lesson contains a wrong figure, a mispronounced name, or a sentence that needs re-wording, and the presenter cannot return to the studio.
Steps. Identify the exact sentence to change and its timecode range. Re-record the sentence, or generate it in the presenter's cloned voice, and cut the new audio to the same duration and cadence as the original where possible. Generate the lip-sync pass over the affected range only, so you pay per second for the fix rather than for the whole lesson. Re-assemble and check the seam at both ends of the replaced segment.
VERIFICATION CHECKLIST for Word-level correction:
[ ] Multi-Model Check: generate the corrected range on two models and compare the jaw and cheek at the seam, where synthetic mouth generation is likeliest to show.
[ ] External Source: watch the corrected range at full speed and then at half speed on a large screen. Check the transition frames at each end, not the middle of the segment, because the middle is where every model does well.
[ ] Human Review: whoever owns the curriculum confirms that the corrected sentence says what it is supposed to say and that the correction is disclosed in the version notes.
[ ] CI-First Test: could the team have made this correction without the tool, and if so, at what cost? If the answer is that the tool replaced a two-minute audio edit, the tool earned its place; if it replaced nothing, the tool added a failure mode.Workflow 3: Producing multi-language variants of a recorded message at volume
The job: one recorded message from a named person needs to go out in four languages to four audiences.
Steps. Confirm consent and rights for the original and for each re-voiced version. Write the translations as scripts rather than as literal translations of the original audio, because cadence and length drive what the mouth has to do. Generate the audio per language, then the lip-sync pass with active speaker detection enabled if more than one person is in frame. Use the batch endpoint if you are on Scale or above and the volume justifies it, since the batch path is what makes a four-language run a repeatable operation rather than four manual jobs. Store the outputs with the language, the model and the run date in the filename or the metadata, so that a later audit can tell what was made and how.
VERIFICATION CHECKLIST for Volume variants:
[ ] Multi-Model Check: spot-check one language on a second model, and spot-check the same timestamps across every language. Errors in this class of work cluster at the same moments across languages, so a spot-check at aligned timecodes catches more than a random sample.
[ ] External Source: a speaker of each target language reviews the version before release. Do not ship a language nobody in the team can evaluate.
[ ] Human Review: a single named owner signs off every language version, and that owner holds the consent records.
[ ] CI-First Test: can the team state, per language, which model ran, what it cost and who approved it? If not, the volume outran the process.Two cross-workflow rules apply to all three.
Never run this on footage of a person who has not consented, whatever the rights position on the footage itself. The tool cannot tell the difference and the terms put the entire burden on you.
Disclose the re-voicing to the audience. The vendor's own product design makes an undisclosed dub look like an original recording, and the provenance watermark that would otherwise reveal it is removed on paid plans. The disclosure has to come from you.
Strengths, Limits, and AI Imposture Risk
Strengths, by dimension
A real research lineage behind the product (Skill and Quality). The team wrote Wav2Lip, released the LSE-D and LSE-C metrics that the literature still uses, and released a benchmark set for this problem. That is a stronger claim on credibility than a product page, and the open-source ancestor is available for anyone who wants to see how the thing works.
A genuine developer surface (Time and Quantity). One endpoint, two official SDKs, webhooks, a documented batch path for up to 500 videos, published concurrency tables, a cost-estimation endpoint and an OpenAPI specification. This is an API-first product with a studio attached rather than a studio with a hidden endpoint, and it shows in the mechanics: asynchronous jobs, terminal statuses, retry guidance and an exportable usage record.
Model choice at a range of prices (Quantity and Quality). Five models from 0.02 to 0.167 per second, with documented differences in architecture rather than only in price. A team can start on the cheap model, prove the workflow, and pay up only where the footage demands it.
A model that solves the pause problem (Quality). The full-shot architecture in sync-3 is the one capability in this family that changes what is possible, because the chunked models fail on still frames and every real edit has still frames.
The engine lives where editors already work (Time). Premiere Pro and DaVinci Resolve plugins and a ComfyUI node put the model inside the existing pipeline, and the changelog records that all surfaces now run on one foundation.
Disclosure advice published by the vendor itself (Critical Thinking, and credit where it is due). The vendor's article on safe AI dubbing says the output can look identical to a fabrication and that the authorization is everything, and advises keeping a record of permission when dubbing a voice that is not yours. A vendor that publishes its own principal risk is doing something most do not.
Limits, by dimension
No independent measurement of lip-sync accuracy on any commercial model (Quality). The vendor's own model comparison table has eight rows that could carry numbers and none of them does. There is no LSE-C figure, no human preference study, no third-party head-to-head with a published method. A buyer is choosing between five prices on the basis of prose.
Internally inconsistent resolution and language claims (Quality). lipsync-2-pro is 512x512 in the documentation model table and 4K-capable on its own product page. The dubbing language count is 29, 92, or 95 or more depending on the surface and the generation of the feature. Neither inconsistency is fatal and both shift the decision.
A per-second price on every attempt (Time and Quantity). Iteration is the normal condition of this work, and every re-run is billed. A workflow that needs three passes costs three times the estimate, and the estimate does not include the human review time.
Segment ceilings and audio ceilings (Time). The generation cap is per job: 20 seconds free, 1 minute on Hobbyist, 5 minutes on Creator, 10 on Growth, 30 on Scale. Audio above 300 seconds is rejected. Long-form work is always assembled.
Chunked inference caveats on the two mid-tier models (Quality). Still frames break lipsync-2 and lipsync-2-pro, and obstruction detection, which fixes the hard cases, slows generation down. The cheap model and the safe model are not the same model.
A narrow single-capability tool (Quantity). It re-times a mouth on footage you already have. It does not write, generate scenes, caption or publish, and it changes the mouth while preserving the face, which is the opposite of what avatar tools do.
Brand and persona governance is absent (Social Authenticity). There is no brand voice layer, no persona approval step, and no output-side attribution. The output is as on-brand as the script and the audio you supply.
Legal documents predating the product (Social Authenticity, Critical Thinking). The terms were last revised on 6 June 2024 and the privacy policy took effect on 27 May 2024. Voice cloning, built-in dubbing, batch dubbing in 92 languages, image-to-video, assistant mode and the MCP server all shipped after those dates. The documents that govern the data have not been refreshed to match the product. Section 7c quotes them.
The provenance signal is switchable by the user (Critical Thinking). The vendor offers a verification path that tells a third party whether a video was modified with its technology, and removes the output watermark on the Creator tier and above. A disclosure mechanism that the operator can turn off for 19 dollars a month is not a disclosure mechanism; it is a product feature.
AI Imposture Risk
Medium overall. One trap is High, two are Medium, and the High trap has a clear mitigation that the vendor's own documentation supports. Evidence follows for each trap.
Trap | Rating | Evidence |
Time Illusion | Medium | The time saving against a reshoot or an ADR session is real and large, and the honest cost inside the tool is not the generation: it is segmentation against the per-generation cap, the per-second billing of every attempt, the documented slower path when obstruction detection is enabled, and a human frame check that no feature removes. A team that budgets for one generation and gets five has lost the saving it came for. Medium rather than Low because the iteration count is not visible before the first run, and Medium rather than High because a fifteen-minute first session on your own footage makes the shape of the cost clear quickly. |
Quantity Illusion | Medium | The tool produces volume that looks finished. Batch mode covers up to 500 videos, and every output is a plausible video of a real person, which means a defect at scale is invisible without frame-level review at scale. There is no published accuracy figure to sample against and no per-output confidence signal, so verification is human and linear while production is programmatic. Medium rather than High because the tool is honest about its input caveats in the documentation and because a named human sign-off per language is a workable control. |
Skill Illusion | High | The product's value proposition is indistinguishable output for users who cannot evaluate output. Assistant mode is documented as recommending the model, detecting the speakers and generating the sync automatically with no settings to configure and no modes to pick, which removes the last place a user would meet the model-selection judgement that decides the result. The studio's own framing is that there is no timeline and nothing to learn. The operator therefore holds an expert-looking post-production capability with no measurement to check it against, and the two documented caveats that break the mid-tier models, still frames and obstructions, are exactly what a non-specialist cannot anticipate. This is the trap the framework's Skill criteria describe directly: expert-looking output for users who lack the skill to evaluate it. |
Overall: Medium. Two traps are Medium and one is High with a clear mitigation, which the framework assigns to the Medium band. The mitigation is not to ban the tool but to structure the work: choose the model on a real test rather than on price or default, keep a named human approver per output, and put the model, the cost and the approver in the record for every version you ship.
Superhuman Usage Guidance
Invite it for language localisation of footage you own and control, for word-level corrections that would otherwise need a reshoot or a cutaway, for building lip sync into a product through the API, and for teaching what synthetic media now does.
Keep it out of any workflow where nobody in the room can evaluate the target language or the footage, any message that is meant to be attributed to a real person without a disclosure, any use on footage of a person who has not consented, and any live or interactive setting, because the product is not built for low latency.
U365 method integration. Run it as a Co-Worker and Assistant (Profile 2) inside a CI-First discipline: attribute the task to the tool before you give it, keep the script and the consent records on the human side, and use the verification checklist per workflow. Pair it with UP-Context so the script and the review standard come from the institution's own voice rather than from the tool.
Over-delegation warning. The failure mode this tool invites is not bad video. It is a team that stops being able to tell good video from bad and stops needing to, because the assistant mode always produces something plausible. Set a rule that a human who can evaluate the target language signs every version, in writing, with the model and the run date recorded next to it. If that rule cannot be met, do not run the batch.
Section 7c: Consent, likeness, and the licence over your footage, stated plainly
This section has a real finding, and it comes from the vendor's own documents rather than from any third party. It is a contract-terms finding and a permissions finding at once, which makes it the most commercially consequential part of this review for anybody in Higher Education. Read it before the first upload rather than after.
The finding, in one sentence. To use the service you must warrant that you own or control the footage and that you have obtained every required consent, and in exchange you grant Sync Labs a perpetual, irrevocable, worldwide, sublicensable licence over that footage, including the right to use it for promotion and marketing, while the same contract states that the vendor does not retain the content you generate longer than necessary to let you use the service.
Clause 1: what you warrant about your own footage
The terms of service, last revised 6 June 2024, place the entire rights and consent burden on the user. The operative text, quoted verbatim:
"You represent and warrant that you own all right, title and interest in and to your User Content, including all copyrights, trademarks and rights of privacy and publicity contained therein, and you have provided any required notices and obtained any required consents to upload, store, transmit, use, or otherwise process User Content via the Service in accordance with applicable laws."
Read that against the product. The product's core function is to re-animate an identifiable person's mouth and, through the dubbing flow, to reproduce an identifiable person's voice. The terms require you to have obtained the consent of the people in the footage and to have given any notices that are required. They do not define what a required notice is, they do not provide a consent form, and they impose no record-keeping obligation on either party. The vendor states the duty; the mechanism is entirely yours.
The terms also prohibit, in their own list of banned uses, uploading content that "infringes any intellectual property or other proprietary rights of any party (including, without limitation, any rights of privacy)" or that "you do not have a right to upload under any law or under contractual or fiduciary relationships", and separately prohibit using the service to "impersonate any person or entity, or falsely state or otherwise misrepresent your affiliation with a person or entity". A deepfake of a named individual is caught by the first prohibition, and the impersonation clause sits on the terms page as a conduct rule rather than as a technical control. Consent records are not required by the contract, which means the practical burden of proving authorisation falls on you at the moment you need it.
This point has extra weight for U365, whose method content is recorded by named instructors, faculty and staff, and whose learners include minors. The vendor's terms require a parent or guardian's express consent for use by anyone under 18 and state that the service is not intended for use by children. Any project involving footage of a U365 Fellow under 18 needs the guardian consent in writing before upload, and the terms put that on you.
Clause 2: the licence you grant
The same document grants the vendor a licence over your footage that is broader than the operation of the service requires. The operative text, quoted verbatim:
"You hereby grant Sync Labs and its affiliates, successors and assigns a non-exclusive, worldwide, royalty-free, fully paid-up, transferable, sublicensable (directly and indirectly through multiple tiers), perpetual, and irrevocable license to copy, display, upload, perform, distribute, store, modify, and otherwise use your User Content, in any form, medium or technology now known or later developed, (a) in connection with the operation of the Service, (b) to develop and improve the Service and other Sync Labs offerings, (c) for the promotion, advertising or marketing of the foregoing; and (d) as otherwise set forth in our Privacy Policy."
Three features of that sentence matter to an institution rather than to a casual user. The licence is perpetual and irrevocable, so withdrawing a video from the service does not withdraw the permission already granted. It is transferable and sublicensable through multiple tiers, so it can travel to parties you never contracted with, including in a corporate transaction, which the same terms confirm by listing business transferees as a category of recipient. And limb (c) is a marketing right: your footage can be used to promote the vendor and its offerings. Limb (b) is the model-improvement right, and a separate usage-data clause authorises the vendor and its third-party service providers to "collect and analyze User Content" and to derive usage data, which the vendor may use "for any purpose in accordance with applicable law and our Privacy Policy".
The privacy policy, effective 27 May 2024, restates the model-training position in its own words. It says the vendor may use personal information for the legitimate interest of improving the service, and that this includes "developing, training, and fine-tuning the AI models we use to support the Services and to develop new features and products". The same policy states that User Content may be uploaded "only as permitted in our Terms of Service".
Why the two documents read against each other are the finding
The terms contain a sentence that, read alone, sounds like a narrow retention promise. Quoted verbatim:
"When you choose to use these features and tools, Sync Labs is acting only on your instructions in order to facilitate the service requested by you and Sync Labs does not retain the User Content you generate through use of such features and tools for longer than necessary to allow you to use the Service."
That sentence is true and it is also narrow. It is about the content you generate through the face-and-voice analysis features, and it says nothing about the content you upload, that is, the original footage, which is what the perpetual licence covers. Read separately, the retention sentence reads like a clean data position and the licence reads like boilerplate. Read together, they answer different questions: the output is not kept beyond what the service needs, while the input has been licensed in perpetuity for improvement, marketing, and onward transfer. A reader who takes only the retention sentence will conclude that nothing is retained. The same privacy policy is explicit elsewhere that the licence travels: it lists business transferees as a category of recipient for "all or any portion of the business or assets".
There is a further asymmetry a reader should see. The privacy policy states that cross-border transfers may occur and that the destination jurisdictions "may not provide the same protections as the data protection laws in your home country", and it states that the service does not respond to Do Not Track signals. For a European institution processing identifiable faces and voices, that is a transfer question rather than a footnote. U365's own legal department owns that analysis; this review states the mechanism and does not perform it.
A contract term that constrains the buyer's product, not the vendor's
One further clause is worth flagging because it is unusual and it limits what a team may build rather than what the vendor may do. Independent comparison coverage of the vendor's terms states that users may not use the API or its output to compete with Sync Labs, a usage restriction worth reviewing before building a competing product on top of it. The restriction is real and it is the kind of term that becomes material the moment a U365 venture or a UIT student project is built on the API. Read it before the architecture is decided, not after.
What the vendor does well here, and it should be said
Two things deserve credit in the same section.
The vendor has published a provenance mechanism. The homepage describes a proprietary watermarking technology with a verification path: a third party can upload a video and check whether it was modified using the vendor's technology. That is a genuine contribution to the disclosure problem in this category. It is also switched off on paid output, which is the finding in the next paragraph rather than a criticism of the mechanism's existence.
The vendor has published its own principal risk. Its article titled "Is AI dubbing safe" states plainly that the same technology which localises a creator's video can be misused "to put words in someone's mouth, and the difference is consent", that a deepfake is the same output produced without agreement, that "the output can look identical", that "the authorization is everything", and that a user should "keep a record of permission whenever you dub a voice that is not yours". It also advises checking whether a provider retains files, trains on content, and deletes on request, and to treat vague answers as a reason to look elsewhere. That is unusually honest vendor writing, and a reader should hold the company to the standard it set for itself: by its own test, the answers above are not vague, they are broad, and the record-keeping duty it recommends is not one its own contract requires.
The watermark and the verification path
The two features are in tension and the tension is the section's practical point. A video that carries the vendor's watermark can be checked by anyone. The vendor's billing documentation states that free and Hobbyist outputs carry a watermark and that Creator and above remove it. Since any professional use starts at the 19 dollar Creator tier, the same tier that makes the tool usable makes its output indistinguishable from an unaltered recording by the vendor's own mechanism. Consent, authorisation and disclosure therefore become entirely the operator's responsibility at exactly the tier where the tool is actually bought. That is the operative governance fact for a college or a studio, and it is why every workflow in Section 10 carries a disclosure step.
One further mechanism is worth naming for anyone building an institutional policy. The product's own dubbing path reproduces an identifiable person's voice from a short clip of them, and its documentation describes this as preserving pitch, tone, cadence and speaking style. The vendor's consent-first claim is that the voice it generates is the voice you uploaded. That is a design intention, not a technical control: nothing in the documented API reference reports whether the voice on the uploaded clip belonged to the person in the uploaded video, or whether either of them consented. The contract says you warrant it. The system does not check it. For a U365 project, the enforceable control is the consent record and the named approver, which is what Section 10's checklists require.
No score changed, and why
Nothing in this section changes any score in Section 13, and the reason is methodological. The CI-First framework measures benefit to the human across Time, Quantity, Quality and Skill, and the Humics; it has no dimension for contract terms, licensing or supplier conduct. The licence, the consent burden and the retention wording are all real findings and none of them is a measurement of what the tool does for the user. Recording them here without touching the score is the correct handling for a finding the framework does not score. A reader who sees an unchanged score next to a governance finding should read the unchanged score as "this did not enter the arithmetic", not as "this was considered and dismissed".
What this section does not do
It does not say the vendor has done anything unlawful, and it alleges no misconduct. It quotes the vendor's own current documents and describes what they say. It is not legal advice and not a data-protection assessment: whether a specific U365 processing activity is lawful under the applicable regime is a question for U365's legal function, not for a research review. It does not recommend avoiding the tool, which would substitute a judgement for the reader's own. It states the terms, names the decision, and leaves the decision where it belongs. And it does not touch the resolution or language inconsistencies in Section 8, which are product-claim questions rather than contract questions and are handled there.
U365 Co-Intelligence Rating
Summary
Dimension | Rating | Note |
CI-First Profile | Primary: Co-Worker and Assistant (2). Secondary: Analyst and Tester (4) | Profile 2 because the tool executes a specified transform on specified inputs. Profile 4 for the cost, model and quality comparison work the tool makes possible and the developer who benchmarks the family |
Collaboration Mode | Centaur | Derived from the framework's own rule in Section 7.2: Imposture Risk Medium or High means Centaur, because Centaur mode is safer. Overall risk here is Medium with one High trap |
CI-First Benefit Score | 5.8 / 10 | CI-First Positive |
Time Benefit | 6 | Large savings against a reshoot or an ADR session on the common case, reduced by segmentation against the generation cap, by iteration billed per second, and by a human frame check that the tool does not do |
Quantity Benefit | 7 | Volume is the tool's strongest dimension: one API call per video, batch to 500 on Scale, SDKs, webhooks, and five models at different prices to match the volume to the footage |
Quality Benefit | 5 | Capped by the total absence of independent measurement of lip-sync accuracy on any commercial model, and by two live internal inconsistencies (resolution on lipsync-2-pro, language count in dubbing). Not scored on measured weakness because no third party has published a measurement |
Skill Benefit | 5 | Scored conservatively per the framework's rule. The API, the documented model differences and the input caveats do teach a real discipline. Assistant mode and the "nothing to learn" framing point the same product at the opposite outcome, and a non-specialist cannot evaluate the output while the output's whole purpose is to be indistinguishable |
Humics Protection Score | minus 1 | One protected, two eroded |
Humics Badge | Humics-Neutral | See the dimension table below |
AI Imposture Risk | Medium (Skill Illusion High, Time Medium, Quantity Medium) | See Section 11 |
Status | Active | Last tested 2026-09-25 |
CI-First Benefit Score
(6 + 7 + 5 + 5) divided by 4 = 5.75, rounded to 5.8.
5.8 is CI-First Positive. That is a recommendation with discipline attached, not a lukewarm verdict. The quantity dimension is genuinely strong, the time saving is real and large on the specific case the tool is bought for, and the honest counterweight is that quality cannot be verified against anything the vendor or anyone else has published.
Why Time is 6 and not higher. The saving against a reshoot is real for a course module update, and the vendor publishes a cost estimate before the job, which is more than most tools in this space do. The cost inside the tool is that the human work does not disappear, it changes shape: segmentation, per-second re-runs, a model choice the buyer must make on prose rather than evidence, and a frame-level review of an output designed to defeat casual review. A team that expects one generation and gets four has paid for the saving twice.
Why Quantity is 7 and not higher. The volume mechanics are excellent and documented. The gap is that volume arrives already looking finished, with no confidence signal per output and no published accuracy to sample against, so verification scales linearly while production scales programmatically.
Why Quality is 5. This is the pivotal number and it is driven by measured absence rather than measured weakness. There is no LSE-C or LSE-D figure for any commercial model, no human preference study, no independent head-to-head with a stated methodology. The vendor's own model comparison table leaves eight measurable rows empty. On top of that the vendor publishes 512x512 in one place and 4K in another for the same model, and 29, 92 and 95 or more languages for the same feature. A buyer cannot select between five prices on evidence.
Why Skill is 5. The framework says to be conservative here, and this tool cuts both ways in an unusually clean line. The documented model differences, the input caveats and the API teach a genuine discipline about what lip sync can and cannot do, and a team that reads the documentation learns something transferable. Assistant mode, framed as "no settings to configure, no modes to pick", and a studio described as "no timeline, nothing to learn", remove exactly that encounter. The same product can build the understanding or skip it, and the default path skips it.
Framework v1.2 clause note
All three pedagogical clauses in framework v1.2 were assessed explicitly. Two return a null and one applies, and the nulls are findings rather than omissions.
5.2.3-a, agent-authored procedural memory. Null. The clause sets a Skill Illusion floor of no lower than Medium for a tool that writes procedural memory on the user's behalf. This tool writes no procedural memory for the user. It produces video files and usage records. It stores assets, generations and projects in the account, and those are outputs rather than memory that later sessions read as instructions. The MCP server is the closest thing to an agent surface, and it exposes the vendor's own operations to a coding assistant rather than letting the assistant author the user's standing instructions. Nothing durable is written about how the user should work, so the clause's floor is not triggered. The Skill Illusion rating of High in Section 11 rests on the Skill Illusion trap's own criteria, which the tool's assistant mode and studio framing meet directly, and not on this clause.
4.2-a, agent-mediated conversation. Null, and the null needs a line of reasoning because the tool does touch a person's likeness. The clause erodes Social Authenticity when agent-authored text is presented as the person's own voice in a human-facing channel, or when agent interaction substitutes for human contact. This tool generates no conversational text and mediates no conversation. It re-times a mouth on footage the user supplies and reproduces a voice the user supplies. The Social Authenticity erosion recorded in Section 13 comes from something else entirely: the absence of any brand or persona governance layer, and the removal of the vendor's own provenance watermark on the tier a professional team has to buy. The clause does not reach either of those, so the null stands and the Humics badge does not rest on it.
7.5, team-level rooms. Null. The clause applies to a shared channel where more than one agent acts, and requires Centaur for any such room. This tool is a single-execution service: one request, one asynchronous job, one output, with no multi-agent room and no shared channel in which several agents act on one another's work. The batch endpoint is many independent jobs, not several agents in a room. The clause does not apply, and Centaur is recorded in Section 13 on the framework's own Imposture Risk rule rather than on this clause.
Humics Protection
Humic | Rating | Reasoning |
Creativity | plus 1 | The tool preserves the user's creative material and adds a capability the user did not have. It regenerates the mouth inside a performance the user shot and directed, and the dubbing flow lets one performance reach several language markets. The creative work, the script, the staging and the performance remain the user's. Credit also goes to the vendor for publishing the limits and the caveats rather than hiding them |
Critical Thinking | minus 1 | Three mechanisms erode it. Assistant mode is documented as choosing the model, detecting speakers and generating with no settings to pick, which removes the one decision, model selection, that determines the result. The output is designed to be indistinguishable from an unaltered recording, so the tool trains the user out of the habit of checking. And the vendor's own provenance mechanism, which would let anyone verify a modification, is removed on the tier any professional team buys |
Social Authenticity | minus 1 | The product reproduces an identifiable person's voice from a clip and re-times that person's mouth to new words, which is the highest-stakes synthetic-media capability in the market for a college. The vendor publishes good consent guidance and requires the user to warrant consent, and the contract asks for no records, provides no consent mechanism, and grants the vendor a perpetual marketing licence over the footage. There is no brand or persona governance layer in the product itself, and the risk of an undisclosed re-voiced message that looks like an original recording rests entirely with the operator. The minus is about the absence of controls, not about the tool's existence |
Humics Protection Score: 1 minus 1 minus 1 = minus 1. Badge: Humics-Neutral.
The band matters here and should be read correctly. At minus 1 the badge is Humics-Neutral, the same band as minus 1 to plus 1, which the framework describes as mixed or balanced effects. This is not a Humics-Risky badge and this review does not claim one. What holds the badge in Neutral rather than Risky is the plus 1 on Creativity: the tool augments a creative process rather than replacing one, it preserves the user's own performance rather than generating a substitute, and the vendor writes publicly and honestly about consent. What keeps it out of Humics-Friendly is that the two eroded Humics are the two that matter most in this category, and neither erosion is fixed by a product feature the user can switch on.
A reader who has seen the sibling review in this series should note the difference in evidence rather than the difference in badge. ElevenLabs carries a Humics-Risky badge in this series on its dubbing and cloning surface. Sync Labs supplies at least part of the engine under one ElevenLabs video surface and sells a directly comparable capability on its own account, so the exposure is real and overlapping. It is nonetheless lighter here on the vendor's own record: Sync Labs published a consent-first article that names the risk in its own words, published provenance verification, and wrote its consent duty into the contract as a user warranty. ElevenLabs' badge rests on its own evidence and this review does not inherit or copy it. On the evidence above, minus 1 with a Humics-Neutral badge is where this tool lands. If the terms are not refreshed and the watermark switch is not reconsidered, a future re-score has a reasonable case for Risky, and that is recorded here as a trigger rather than as a current rating.
Superhuman usage guidance, condensed
Use it as a Co-Worker and Assistant. Give it a specified mouth transform on footage you own. Keep the script, the consent and the disclosure on the human side. Choose the model on a real test of your own footage rather than on price or default, because the two documented caveats, still frames and obstructions, decide which model you can use. Put the model, the cost and the named approver in the record for every version. Review the output at frame level at the seams, not at the middle. And if the team cannot evaluate the target language or the footage, do not run the job.
What Users Say
Independent review coverage of this product is thin, and the thinness is itself the finding. There is real signal, and it is mostly from directories that publish an editorial score with no visible user reviews behind it, plus one sentiment aggregator. Numbers below are recorded exactly as published, with the platform and the date where the platform gives one.
Platform | Signal | Notes |
G2 | No reviews found on G2 | Checked directly |
Capterra | No reviews found on Capterra | Checked directly |
Trustpilot | No reviews found on Trustpilot | Checked directly |
GetApp | No reviews found on GetApp | Checked directly |
Product Hunt, sync-3 launch | 132 upvotes, 13 comments, 11th on the daily leaderboard, 7 April 2026 | A second directory records the same launch at 135 upvotes. Read the two as the same event counted differently |
Product Hunt, Sync v2 launch | 110 upvotes, 1 comment, 21st on the daily leaderboard, 24 April 2025 | The earlier of the two launches, on the previous model generation |
RightAIChoice sentiment scan | 45 mentions across 3 sources, 13 per cent positive, researched 27 July 2026 | Sources given as YouTube, Bluesky and Lemmy. This is the only sentiment aggregation found, and it is a small base |
SwitchTools directory | 4.5 out of 5 overall score, 0 user reviews, updated 18 June 2026 | The directory states the review count as zero while publishing a score, which is the point to keep in view |
PopularAiTools.ai directory | 4.1 out of 5, described as rated 4.14 out of 5 by users | No user review count and no individual reviews are shown behind the figure |
Kompozy editorial review | 4.2 out of 5 overall, 3.8 out of 5 on pricing transparency and value | An editorial score, not a user review |
ToolMango editorial review | No score published. Cons listed as narrow single-purpose scope, extreme head angles and occlusions still breaking sync, and costs rising with length and resolution | An editorial summary |
No thread-level figure obtained | No number in this section is attributed to this platform | |
YouTube | Multiple independent comparisons exist | Three of them are the sources for this section and are embedded in the recommendations section below |
What users and reviewers consistently praise, across the sources above. Complex-scene handling, specifically objects crossing or covering the face, which the vendor's sync-3 model is built for and which reviewers credit. The output specification: 4K and ProRes at 60 frames per second, which is the specification a professional pipeline needs. The integrations: Premiere Pro, DaVinci Resolve and ComfyUI named repeatedly as the reason the tool fits a real edit pipeline. Output quality described as impressive. And the targeting itself: reviewers praise that the tool is aimed at professional post-production rather than casual use.
What users and reviewers consistently criticise. Price, and specifically the price model rather than the price level: the subscription plus per-second usage structure is described by more than one source as hard to predict at volume, and one source frames it as the biggest barrier to adoption. The absence of community discussion, which two aggregators state plainly, with one recording discussion volume as low and trending down. The learning curve, described as intermediate and a few hours, with users reported as getting stuck on the sync model parameters and on setting up the editing integrations. Scope narrowness, with the tool described as a single-capability product that re-syncs a mouth and does nothing else. And remaining failure modes on extreme angles and occlusions, reported by one editorial source even though the vendor's flagship model is explicitly sold on handling them.
Three discrepancies in the third-party coverage are worth naming, because a reader who encounters them will otherwise think the vendor moved.
The first is pricing. A third-party review states the Growth tier at 99 dollars a month. The vendor's own billing documentation states 49 dollars a month for Growth. The vendor's figure is the one the service follows, and the third-party figure is either stale or wrong. Treat any third-party plan price here as provisional.
The second is the free tier. One aggregator states that Sync Labs has "no free tier beyond basic trial" and that whether the product is freemium is unclear. The vendor documents three free generations a month with no card required, 10 text-to-speech generations, and all lipsync family models available including lipsync-2-pro. Both statements can be true of the same product if the aggregator counted only what a free user can produce for a real project rather than for a test. The vendor's documentation is the accurate description of the allowance and the aggregator's is the accurate description of the practical value of that allowance.
The third is DaVinci Resolve. One directory lists limited integration as a con and states that native DaVinci Resolve support is not available, requiring export and re-import. The vendor's documentation carries a DaVinci Resolve Studio installation guide and the changelog records dubbing across Adobe Premiere Pro and DaVinci Resolve. The vendor's documentation is current and the directory entry is not.
One credit is due in this section. The vendor is not the source of the negative sentiment: the price criticism comes from users and reviewers, and the vendor publishes a cost-estimation endpoint and an in-app estimate precisely so a buyer can see the bill first. What the vendor does not publish is any measurement of lip-sync accuracy, which Section 11 scores.
U365 editorial note
The signal quality here is low, and no number in this section should be treated as a measurement. Two directories publish scores with zero visible user reviews behind them, one publishes a decimal rating attributed to users without showing the users, and the only sentiment aggregation rests on 45 mentions across three platforms. That is enough to say what people talk about and not enough to say how good the product is. Where a directory score and the vendor's documentation disagree, this review follows the vendor's documentation for pricing and capability and the directories for sentiment.
Comparison and Alternatives
Five alternatives are worth naming, and they split by what they do to your footage rather than by quality. The prices below are published list rates, not negotiated volume rates, and every vendor in this category discounts at contract scale.
Tool | What it does | Published rate | Choose it if |
sync. (Sync Labs) | Re-times the mouth on footage you already have, at up to 4K, with a full-shot model that handles pauses, obstructions and close-ups. API, SDKs, two plugins, MCP server | From 0.04 per second on lipsync-2 to 0.133 on sync-3, plus a 5 to 249 dollar subscription | You have the footage and the rights, you need the mouth to match new audio, and you want a documented API with batch and webhooks |
HeyGen | Dubbing with lip sync on real footage, and separately AI avatar generation and text-to-video. Wider product, 175 or more languages claimed | 0.0333 per second in speed mode and 0.0667 per second for precision lip sync, pay as you go | You want avatar or text-to-video generation alongside dubbing, or you want a single vendor for both, and you accept that the presenter may be generated rather than recorded |
Hedra | Character video with strong performance on stylized and animated inputs, billed against a plan | Character-3 at 6 credits per second against plans from 15 to 75 dollars. Polling-only in the public documentation with no documented bulk endpoint | Your input is stylized or animated and quality per dollar matters more than batch throughput or latency predictability |
ElevenLabs video workspace | A broader creative interface that exposes several underlying models, including the vendor's sync-3 and lipsync-2-pro, alongside voice generation and cloning | Bundled into the platform's own subscription and credit model | You are already on the platform for voice and want lip sync in the same place, and you accept one more layer between you and the model |
Rask AI and similar dubbing suites | End-to-end video translation and dubbing with lip sync, sold as a finished workflow rather than a primitive. 130 or more languages claimed | Plan-based, not published per second | You want a translated video as the product rather than an API call to build on, and you will trade control for convenience |
Open-source line: Wav2Lip and its successors | Self-hosted lip sync. Wav2Lip is the vendor's own research, published in 2020, with the code and the benchmark set released | Free, less your compute | You need full data control, no third-party licence over the footage, and no per-second cost, and you have the GPU capacity and the engineering time |
One structural point about the table. The vendor's own comparison content lists its language coverage below three of these competitors: a vendor article puts sync. at 95 or more languages against Rask AI at 130 or more, Synthesia at 140 or more and HeyGen at 175 or more, and states that the competitor counts are taken from each tool's own site. The vendor's own headline claim on lip sync is quality and control rather than language breadth, and its table reflects that honestly.
The self-hosting question deserves its own line because of what this vendor specifically is. The company's founders wrote Wav2Lip, released the model, the code and the ReSyncED benchmark, and published the metrics. A team with a strict data-residency requirement, such as a public institution handling footage of minors, can run the ancestor of this exact technology on its own hardware and grant no third-party licence over the footage at all. It will not match the current commercial models on hard shots, and the vendor's own blog is the source of the caveats that explain why. That trade is a legitimate institutional answer to Section 7c.
Choose sync. if the mouth must match existing audio on footage you control and the work has to run programmatically. Choose HeyGen if you also need the presenter to be generated. Choose an end-to-end dubbing suite if you want the finished video rather than the primitive. Choose the open-source line if the footage cannot legally or ethically leave your infrastructure, and accept the quality gap.
Verdict and Next Steps
sync. is the most capable dedicated lip-sync engine reviewed in this series and it scores 5.8 out of 10, CI-First Positive. The score and the verdict point the same way with different emphases: quantity and the time saving on the case the tool is bought for are genuinely strong, quality cannot be verified against anything anyone has published, and Skill is held at 5 because the same product both teaches a real discipline through its documentation and, through assistant mode, offers to skip it.
What earns the recommendation. A real research lineage with published metrics. An API-first product with two official SDKs, documented rate limits and concurrency, a batch path for up to 500 videos, a cost-estimation endpoint and an OpenAPI specification. Five models at prices from 0.02 to 0.167 per second with differences in architecture rather than only in price. A flagship model whose full-shot design solves the pause and still-frame problem that breaks the cheaper models, which matters because every real edit has pauses. Plugins inside Premiere Pro, DaVinci Resolve and ComfyUI. A fifteen-minute path from sign-up to a first result with no card required. And a vendor willing to publish, in its own words, that the technology can put words in someone's mouth and that authorization is the whole difference.
What holds the score down. No independent measurement of lip-sync accuracy for any commercial model, against a comparison table with eight empty measurable rows. Two live internal inconsistencies in the vendor's own surfaces, one on the resolution of the mid-tier model and one on the dubbing language count. A per-second charge on every attempt, on a workflow whose normal condition is iteration. The provenance watermark removed on the tier a professional team has to buy. Legal documents dated 2024 that predate voice cloning, bulk dubbing, image-to-video and assistant mode, still granting the vendor a perpetual marketing licence over uploaded footage of identifiable people while requiring no consent record from anyone. And a Skill Illusion trap rated High, which is why the framework puts this tool in Centaur mode.
The recommendation to a U365 reader, in one paragraph. Adopt it for language localisation and word-level correction on footage U365 owns, run it in Centaur mode with a named human approver per version and a written record of the model, the cost and the approver, choose the model on a test of your own footage rather than on price or default, and write the consent and disclosure rule before the first upload rather than after. Do not run it on footage of a person who has not consented, whatever the rights position on the footage itself.
Next steps
A fifteen-minute evaluation, in order. Create a free account, no card. Take one clean face-to-camera clip you own, 15 to 20 seconds. Run the same segment on lipsync-2 and on sync-3 and watch both at full speed and at half speed at the seams. That single comparison answers the model question that the vendor's documentation leaves qualitative. Then run the cost-estimation endpoint for a realistic job and compare it against your reshoot or ADR baseline. If both comparisons pass, subscribe at Creator for the watermark-free tier and run the first real job with the Section 10 checklist attached.
A governance step that runs in parallel and does not depend on the test. Before any institutional upload, U365's legal function should read the two clauses quoted in Section 7c and decide the institution's position on the perpetual marketing licence, the sublicensing, the cross-border transfer statement, and the consent record the contract asks for but does not define. That decision is upstream of every workflow here.
Status and Last Tested
Status: Active | Last tested: 2026-09-25 | Re-check: a new lipsync or react model release, a published lip-sync accuracy measurement from any source, a revision of the terms of service or the privacy policy, a change to which plans carry the output watermark, or a reconciliation of the dubbing language count.
Active: the tool is current and recommended.
For U365 deployment
Run it as a Co-Worker and Assistant (Profile 2) inside a CI-First discipline, with UP-Context supplying the script and the review standard so the institution's voice is not delegated to the tool. Attribute the task to the tool before you give it, keep the consent record and the editorial judgement on the human side, and apply the Executive Safeguard to every output: assume the worst plausible result and check for it at the seams rather than in the middle. Two U365-internal conditions apply. The named approver must be someone who can evaluate the target language, not merely someone available. And any version made with this tool must carry its disclosure in the same metadata as the version itself, so an audit does not depend on anyone's memory.
UP-Context prompt pack
Five prompts to run through UP-Context before an institutional deployment. Each one carries a Role line, written in the U365 prompting method, and a verification close that states what the human does with the output.
Prompt pack 1: The consent record, written before the first upload
Role: AI as a media-rights and consent-record drafter for a higher-education publisher.
Given U365's publishing and teaching obligations, draft the consent record template our presenters will sign before any footage of them enters a lip-sync or voice-cloning workflow, naming the licence we grant upstream, the retention we require, and the disclosure we will attach to every output.
UP-Context verification: state what you checked, what you could not check, and what a human must decide before this is used.Prompt pack 2: The sign-off record, with the model, the cost and the approver named
Role: AI as a quality-assurance officer writing a sign-off record for an editorial process.
Our rule is that no lip-synced or re-voiced version reaches a learner without a named human who can evaluate the target language. Write the sign-off record for that rule, including the model used, the run date, the per-second cost and the disclosure text.
UP-Context verification: state what you checked, what you could not check, and what a human must decide before this is used.Prompt pack 3: The policy review against CI-First, with the cross-vendor calibration
Role: AI as a co-intelligence reviewer auditing a draft policy against a published framework.
Review our draft policy against the CI-First discipline: where in this workflow are we delegating a judgement that belongs to a human, and which of the three Imposture Traps are we leaving unmitigated? Also name the second engine, from a different laboratory, that will be run on one segment, and record the per-seam disagreement count next to the model, the cost and the named approver. A comparison between two models from the same vendor tests the vendor's own family and does not test the vendor's claim.
UP-Context verification: state what you checked, what you could not check, and what a human must decide before this is used.Prompt pack 4: The learner-facing explanation of a re-voiced version
Role: AI as a plain-language editor writing for adult learners.
Write the learner-facing explanation of why a U365 module is available in several languages and how the language version was produced, in the institution's own voice, without overstating what the technology does.
UP-Context verification: state what you checked, what you could not check, and what a human must decide before this is used.Prompt pack 5: The instructor note on what changes and what does not
Role: AI as an internal-communications writer addressing teaching staff.
Draft the internal note that tells instructors what changes and what does not when their recorded module is localised, so that the same performance in a new language is understood as a re-voicing rather than a re-recording.
UP-Context verification: state what you checked, what you could not check, and what a human must decide before this is used.Migration Path
Not applicable. sync. (Sync Labs) is an Active tool under this framework: it is current, it is recommended, it has shipped new models and new surfaces within the review window, and nothing in this review identifies a competitor that has clearly surpassed it on the lip-sync capability itself. No migration plan is warranted. If the consent, licensing and provenance findings in the section before this one are not addressed at a future re-score, that assessment could move to Risky, at which point this section would carry the open-source line from the comparison as the recommended replacement, the data-control advantage as what transfers, the hard-shot quality gap as what does not, and self-hosting capability as the migration prerequisite.
U365's Recommendations to Learn More
The vendor's documentation is where the caveats live, and it repays the time more than the marketing pages do.
Official learning resources
The documentation index, the machine-readable starting point for every surface: https://sync.so/docs/llms.txt
The product introduction and the at-a-glance specification block: https://sync.so/docs/introduction
The quickstart, which carries the SDK installation, the input limits and the vendor's own pitfall list: https://sync.so/docs/quickstart
The model family reference, including the comparison table and the caveat list that appears nowhere else in the vendor's public material: https://sync.so/docs/models/lipsync
The flagship model reference and the free-account limit: https://sync.so/docs/models/sync-3
The billing documentation, which explains the subscription plus usage model, the invoice thresholds and the per-output-frame charge: https://sync.so/docs/product/billing
The free-trial allowance: https://sync.so/docs/free-trial
The API reference for rate limits, concurrency and retry guidance: https://sync.so/docs/api-reference/guides/rate-limits
The release history, including dubbing to 92 languages: https://sync.so/changelog
Video tutorials and channels
The three walkthroughs below were each resolved against the YouTube oEmbed endpoint, and each is an independent explanation rather than a vendor demonstration. The first is a four-tool comparison that reads out this vendor's model list, its plan pricing and its per-second rates, and it is useful as a cross-check on the pricing table in this review. The second is a broad round-up that includes a segment generated with this vendor's technology and states the free-tier terms. The third is an early hands-on walkthrough from before the current model generation, and it shows how far the output has moved.
Written tutorials and deep-dive articles
The vendor's article on whether AI dubbing is safe, which is its own statement that the output can look identical whether or not the person agreed and that authorization is the whole difference: https://sync.so/blog/is-ai-dubbing-safe
The vendor's one-pass dubbing workflow for YouTube videos, which is where the competitor language table quoted in this review appears: https://sync.so/blog/how-to-translate-youtube-videos
The vendor's account of the Wav2Lip lineage, its citations and its benchmarks: https://sync.so/blog/what-is-ai-lip-sync
The vendor's account of the LSE-D and LSE-C metrics and the ReSyncED benchmark: https://sync.so/blog/a-lip-sync-expert-is-all-you-need-for-speech-to-lip-generation-in-the-wild
The Wav2Lip paper itself, the source for the two metrics the field still uses: https://purehost.bath.ac.uk/ws/files/211225619/Wav2Lip_ACMMM_20_camera_ready_2_.pdf
Community and social
The vendor's Discord community, referenced from the documentation and the fastest surface for release news
The model card for a variant of this family on a third-party marketplace, useful when you want to compare a hosted rate against a self-hosted one: https://replicate.com/sync/lipsync-2
The practitioner post quoted in this review, in which the same class of tool was used on a live production surface, reached at https://x.com/joecolantonio
Resources on X
Dedicated X channels:
The account to add first is the vendor's own at https://x.com/synclabs, because a change to the pricing model, the licence over your content or the model line would be announced there before it reached a documentation page.
Treat the vendor's own channel as a promotional source. It is useful for learning that a release happened, and not for assessing whether an accuracy claim holds.
The two cautions stated above apply to every third-party video in this section. One comparison reads the legacy model's rate as roughly 2 cents per second, which matches the vendor's documentation, and it states that a free sync-3 generation is capped at 15 seconds, which also matches. Where a video and the vendor's documentation disagree, this review follows the documentation. And the oldest of the three describes a much earlier version of the product with a credit-based free allowance that no longer exists, so treat its pricing as history rather than as a current figure.
Resources on sync. (Sync Labs)
Read the vendor's documentation before its marketing pages, and read three of its articles beside it. The lip-sync model comparison page carries the family table and, more importantly, the caveat list: the still-frame limitation on lipsync-2 and lipsync-2-pro is documented there and nowhere else in the vendor's public material, and it is the single most useful operational fact about this family. The billing page explains the subscription plus usage model, the invoice thresholds, and the fact that rates are quoted at 25 frames per second while billing follows output frames, and a reader who skips it will misprice the first real job. And the article on whether AI dubbing is safe is the vendor's own statement that the output can look identical whether or not the person agreed; read it beside the terms of service, with the consent and licence section of this review between them. https://sync.so/docs/models/lipsync, https://sync.so/docs/product/billing, https://sync.so/blog/is-ai-dubbing-safe
Glossary
ADR
Automated Dialogue Replacement. The traditional post-production process of re-recording an actor's dialogue in a studio and fitting it to the picture. It is the workflow this class of tool is priced against.
CI-First
University 365's core method for structuring work between human intelligence and artificial intelligence. CI comes from HI plus the product of AI and HI, which is why a tool that erodes the human term reduces the total even when the machine term rises.
CI-First Benefit Score
The single headline number in a U365 Tools Review, computed as the average of four dimensions scored from 0 to 10: Time, Quantity, Quality and Skill. Bands: 0 to 2 CI-First Negative, 2.1 to 4.0 CI-First Neutral, 4.1 to 6.0 CI-First Positive, 6.1 to 8.0 CI-First Strong, 8.1 to 10 CI-First Transformative.
CI-First Profile
One of five roles attributed to an AI tool before it is given work: Co-Creator and Thought Partner (1), Co-Worker and Assistant (2), Coach and Tutor (3), Analyst and Tester (4), Challenger and Devil's Advocate (5). Attribution happens before the task, not after.
Centaur
A Collaboration Mode with a clear division of labour. The human handles the Humic work, meaning judgement, taste, empathy and final approval, and delegates the heavy processing to the tool. Framework v1.2 requires Centaur where a tool's AI Imposture Risk is Medium or High.
Collaboration Mode
The structural choice of how human and tool work together across a whole task, either Centaur or Cyborg.
Concurrency
The number of generations one account may have in progress at the same time. On this product: 1 on Free and Hobbyist, 3 on Creator, 6 on Growth, 15 on Scale.
Cyborg
A Collaboration Mode of continuous, intertwined co-creation with rapid iteration in real time. It carries the highest over-delegation risk and is reserved for tools whose AI Imposture Risk is Low.
Executive Safeguard
The U365 discipline of assuming you are working with the worst available AI and building verification accordingly, rather than trusting the tool's own confidence in its output.
Humics
The set of uniquely human capabilities, from Pascal Bornet's work: creativity, critical thinking and social authenticity. Each is scored plus 1, 0 or minus 1 in this framework.
Humics Protection Badge
The label derived from the sum of the three Humics ratings: plus 2 to plus 3 Humics-Friendly, minus 1 to plus 1 Humics-Neutral, minus 2 to minus 3 Humics-Risky.
LSE-C and LSE-D
Lip-Sync Error Confidence and Lip-Sync Error Distance. Two metrics introduced by the Wav2Lip paper and still used across the lip-sync literature to score how well generated mouth movement matches audio. No vendor in this category publishes either against its commercial models.
MCP
Model Context Protocol, a standard that lets an AI assistant call external tools. The vendor ships an MCP server so a coding assistant or chat client can drive generations directly.
Per-output-frame billing
The vendor's usage model. Rates are quoted as a per-second figure measured at 25 frames per second, but the charge is calculated on the frame count of the video you produce, so a higher frame rate master costs more than the per-second figure suggests.
Provenance watermark
A signal embedded in an output that lets a third party verify that a video was modified using a given technology. This vendor offers a verification path and removes its watermark on the Creator tier and above.
ReSyncED
A benchmark set of real synced videos released with the Wav2Lip paper, created because the benchmarks in use at the time could not distinguish a well-synced model from a badly synced one on real footage.
Review Status
The standing of a tool at the date of review, using the U365 vocabulary: Active, meaning current and recommended; Changed, meaning a material update has landed and a re-score is pending; Risky, meaning the tool has significant unresolved issues or has been clearly surpassed and should be used with caution; Deprecated, meaning superseded and in wind-down; Retired, meaning withdrawn.
Visual dubbing
Changing the language of a video while re-timing the speaker's mouth to the new audio, so the result looks recorded in the target language rather than dubbed over it.
Webhook
A callback the vendor sends to a URL you control when a job changes state, so a pipeline does not have to poll for the result.
Zero-shot
Running a model on a person it has never seen, with no per-person training or fine-tuning. It is the property that makes this class of tool practical, because it means one upload is enough.
User Sentiment
The aggregated public opinion from review platforms, community forums and repository activity, reported separately from the CI-First score because crowd sentiment can disagree with a rigorous evaluation. Where the two agree the finding is stronger, and where they diverge the divergence is worth explaining. sync. (Sync Labs) has no reviews on G2, Capterra, Trustpilot or GetApp, and the only sentiment aggregation found rests on 45 mentions across three platforms, so the signal is thin and directional rather than measured.
Sources
Every source below was used in this review and each resolved when checked on 2026-09-25.
Vendor primary sources.
https://sync.so/ product homepage, model claims and the provenance verification feature
https://sync.so/about company, founding year, investors and location
https://sync.so/pricing pricing page
https://sync.so/changelog release history, including dubbing to 92 languages
https://sync.so/changelog/dubbing-assistant-mode-mcp-server-and-sync-3 the dubbing, assistant mode, MCP server and sync-3 release
https://sync.so/changelog/clearer-workflows-better-dubbing-and-a-smoother-studio the newer dubbing engine and studio changes
https://sync.so/terms terms of service, last revised 6 June 2024
https://sync.so/privacy privacy policy, effective 27 May 2024
https://sync.so/docs/introduction the product introduction and at-a-glance specification block
https://sync.so/docs/models/lipsync the lip-sync model family, comparison table and caveat list
https://sync.so/docs/models/sync-3 the flagship model reference and free-account limit
https://sync.so/docs/product/billing plans, usage billing, invoice thresholds, refund policy and watermark rules
https://sync.so/docs/free-trial the free-tier allowance
https://sync.so/docs/api-reference/guides/rate-limits rate limits, concurrency and retry guidance
https://sync.so/docs/quickstart the SDK quickstart, input limits and pitfalls
https://sync.so/docs/llms.txt the documentation index
https://sync.so/lipsync-2-pro the mid-tier model product page and its 4K claim
https://sync.so/sync-3 the flagship model product page
https://sync.so/blog/is-ai-dubbing-safe the vendor's own dubbing safety and consent guidance
https://sync.so/blog/how-to-translate-youtube-videos the one-pass dubbing workflow and the vendor's competitor language table
https://sync.so/blog/what-is-ai-lip-sync the vendor's account of the Wav2Lip lineage and its GitHub star and citation counts
https://sync.so/blog/a-lip-sync-expert-is-all-you-need-for-speech-to-lip-generation-in-the-wild the vendor's account of the LSE-D and LSE-C metrics and the ReSyncED benchmark
Independent and third-party sources.
https://replicate.com/sync/lipsync-2 the model card for lipsync-2 on a third-party model marketplace
https://hunted.space/product/sync-9 the sync-3 Product Hunt launch, recorded at 132 upvotes and 13 comments
https://hunted.space/product/sync-v2 the Sync v2 Product Hunt launch, recorded at 110 upvotes and 1 comment
https://vibecrowd.ai/launches/sync the vendor's launch profile, recording a Product Hunt figure of 135 upvotes
https://kompozy.io/reviews/sync an editorial review with an overall score of 4.2 out of 5 and a pricing-transparency score of 3.8 out of 5
https://rightaichoice.com/tools/sync-labs a directory profile with pricing, integrations and reported frustrations
https://rightaichoice.com/tools/sync-labs/sentiment the only sentiment aggregation found: 45 mentions, 13 per cent positive, researched 27 July 2026
https://www.switchtools.io/tool/sync a directory profile publishing a 4.5 out of 5 score with zero user reviews
https://popularaitools.ai/tools/sync-so a directory profile with a 4.1 out of 5 rating
https://toolmango.com/tools/sync-labs an editorial summary listing scope narrowness and residual occlusion failures
https://techpilot.ai/tools/sync-so a directory profile on the vendor's origin story
https://versely.studio/blog/best-lip-sync-api-2026 a four-engine API comparison with published list rates
https://hokai.io/hub/tools/sync-labs a directory profile with a plan and rate breakdown
https://elevenlabs.io/video/sync-3 the surface on which the vendor's flagship model is exposed inside another platform's model library
https://purehost.bath.ac.uk/ws/files/211225619/Wav2Lip_ACMMM_20_camera_ready_2_.pdf the Wav2Lip paper itself, for the LSE-D and LSE-C definitions and the ReSyncED benchmark
Video sources, each verified live through the oEmbed endpoint on 2026-09-25.
https://www.youtube.com/watch?v=llz5KhkquyI a four-tool comparison that reads out this vendor's model list, plans and per-second rates
https://www.youtube.com/watch?v=JCRwZZGQbiU a round-up containing a segment generated with this vendor's technology and stating its free-tier terms
https://www.youtube.com/watch?v=Vbcyny7dvEg an early hands-on walkthrough of the product before the current model generation
Faculty Note on Evidence Quality
The evidence behind this review is strong on the vendor's own surfaces and weak on independent measurement, and the weakness is the reason several scores are capped rather than the reason they are high. Four things deserve a reader's attention.
First, no independent measurement of lip-sync accuracy exists for any commercial model in this family, and the vendor publishes none itself. The company's own research introduced LSE-D and LSE-C and released the ReSyncED benchmark precisely because earlier benchmarks could not separate a well-synced model from a badly synced one on real footage. Those metrics are still cited across the literature, including in recent academic work on lip synchronization. None of them is applied to sync-3, lipsync-2-pro, lipsync-2 or react-1 anywhere this review could find. The vendor's model comparison table carries eight rows that could hold a number and every one of them is empty, leaving qualitative prose and a single relative speed figure, 1.5 to 2 times slower than lipsync-2 for lipsync-2-pro. Independent academic work in this field does publish measurable comparisons, including landmark distance and lip-sync error confidence figures against named models. That work does not cover this vendor's commercial models, and the gap between what the field can measure and what the buyer can read is the finding.
Second, the vendor publishes no headline percentage of any kind, so there is no accuracy percentage to audit against its own base. That is the honest reading of a product page that leads with qualitative claims such as nearly indistinguishable rather than with a number. Where the vendor does publish a comparison, its base is stated: a vendor article lists this product's language coverage below three named competitors and states that the competitor counts are taken from each tool's own site. Publishing a lower number than your competitors and naming the source of their numbers is the right behaviour, and it is worth saying so. The one cross-vendor comparison with a weaker-setup risk sits on the other side of the fence: the vendor's flagship model also appears as a selectable option inside ElevenLabs' own video workspace, so a reader comparing the two products may be comparing a platform against an engine it partly depends on. Nothing in either vendor's material resolves which model runs under that surface, and a reader should not assume the comparison is like for like.
Third, where the vendor publishes several figures for one thing, this review names each with its surface and treats the spread as the finding, per the framework's rule. Two spreads are material. The resolution of lipsync-2-pro is 512x512 with enhanced detail preservation in the documentation's model table, 512x512 in the same documentation set's question and answer, and 4K on the model's own product page. The dubbing language count is 29 languages in the release note for the first-generation dubbing flow, 92 languages in the changelog for the current engine, and 95 or more languages in the model documentation and a workflow article. A third spread, the rate limit on the generate endpoint, is 100 requests per minute in the current documentation and 60 requests per minute on an older published page, and it is the one of the three a buyer can settle at runtime. A fourth, the plan price for the Growth tier, is 49 dollars a month in the vendor's own billing documentation against 99 dollars a month in a third-party review; the vendor's figure is the one the service follows.
Fourth, the review-platform signal is thin and partly circular. This product has no reviews on G2, Capterra, Trustpilot or GetApp, which this review records as a finding rather than as an absence to be filled. Two directories publish an overall score, 4.5 and 4.1 out of 5, and one of them states its user review count as zero on the same page as the score. Two Product Hunt launches are recorded, and different directories record the same launch with different upvote counts, 132 and 135, so the counts are directional rather than precise. The only sentiment aggregation found rests on 45 mentions across three platforms. That base is enough to say what people discuss and not enough to say how well the product works. No figure in this review is attributed to a page whose content could not be read, and the platform-level rows above say so rather than carrying a number that no page supports.
Two statements about this review's own limits. Its product pricing, free-tier terms, rate limits, model list and plan tiers come from the vendor's documentation and billing pages, which are more current than any third-party profile checked here, and a third-party figure should be treated as provisional wherever the two disagree. And the legal analysis in Section 7c quotes the vendor's current documents and does not extend beyond them: it is not a legal opinion, not a data-protection assessment, and not advice that any particular U365 processing activity is lawful. That determination belongs to U365's legal function, which should read the two clauses quoted there before any institutional upload.
UNOP alignment. UNOP is University 365 Neuroscience-Oriented Pedagogy, and three notes apply to this tool.
Supports. A recorded module localised into a second language supports multi-modal repetition: the same material met through reading and through listening, in the Fellow's stronger language and in the target language. That is a study pattern the method favours, and the tool makes it affordable on a library that already exists.
Supports. Word-level correction without a reshoot keeps the recording current. A module containing a wrong figure or a mispronounced name is corrected at the sentence rather than withdrawn, which keeps the learner's review material accurate rather than stale.
Conflicts, and this is the controlling note. Listening is a passive channel. The method's principle rests on active recall and spaced retrieval, so a localised audio or video version supplements study of material the Fellow has already worked through and does not substitute for it. Used as a substitute it works against the principle the method relies on.
Recommendation. Adopt it as a production tool and not as a study aid. Run the tool on material already studied, never in place of studying it, and treat a localised version as the same content in another language rather than as new learning.
CI-First Evaluation Summary Card
Field | Value |
Tool | sync. (Sync Labs) |
Status | Active |
Last tested | 2026-09-25 |
CI-First Benefit Score | 5.8 / 10 |
Band | CI-First Positive |
Time Benefit | 6 / 10 |
Quantity Benefit | 7 / 10 |
Quality Benefit | 5 / 10 |
Skill Benefit | 5 / 10 |
CI-First Profile | Primary: Co-Worker and Assistant (2). Secondary: Analyst and Tester (4) |
Collaboration Mode | Centaur |
Humics Protection Score | minus 1 |
Humics Badge | Humics-Neutral |
Creativity | plus 1 |
Critical Thinking | minus 1 |
Social Authenticity | minus 1 |
AI Imposture Risk | Medium |
Time Illusion | Medium |
Quantity Illusion | Medium |
Skill Illusion | High |
Clause 5.2.3-a | Null: the tool writes no procedural memory on the user's behalf |
Clause 4.2-a | Null: the tool generates no conversational text and mediates no conversation |
Clause 7.5 | Null: single-execution service, no shared multi-agent room |
Section 7c | Applies: consent, likeness and the licence over uploaded footage |
Key strength | A full-shot model that handles the pauses and obstructions where the cheaper models fail, on a documented API with batch, webhooks and two SDKs |
Key limit | No independent measurement of lip-sync accuracy for any commercial model, with eight measurable rows left empty in the vendor's own comparison table |
Recommended mode | Centaur, with a named human approver per version and a written record of the model, the cost and the approver |









Comments