top of page
Abstract Shapes

INSIDE

PUBLICATIONS

Manus AI: the autonomous agent that writes its own procedures, scored 5.3 on the U365 CI-First Review

1 day ago
82 min read
Manus AI, the general autonomous agent at manus.im: the vendor's own social card

Status: Active | Last tested: 2026-09-25 (Manus, as documented at manus.im in September 2026) | Re-check: trigger-based (max 6 months)


Active: the tool is current and recommended.


What Active means here. Active means current and recommended for the work this review describes: research-shaped deliverables, many-item processing, and first-pass builds where you read the result before anything depends on it. It does not mean the product is predictable. The credits a task will consume are not disclosed before it runs, the vendor publishes no independent measurement of reliability, and the two review platforms that carry the most weight for a paid software decision rate it poorly on exactly the issues this review is about. A reader who needs a fixed cost per task, or a measured guarantee, should treat those as unavailable for now and plan around them rather than assume they will improve.


Version reviewed: Manus, as documented at manus.im in September 2026.


For detailed explanations of the CI-First evaluation terms used in this review, including the Humics Protection Badge and the AI Imposture Risk levels, see the Glossary at the end of this post.




Manus AI Review
Back to the TOC

In this Tool Review



Back to the TOC

Status and Re-check


Status: Active | Last tested: 2026-09-25 (Manus, as documented at manus.im in September 2026) | Re-check: trigger-based, maximum 6 months


Status and Last Tested


Status: Active Last tested: 2026-09-25 Version reviewed: Manus, as documented at manus.im in September 2026 Next re-test: Trigger-based, maximum six months. The triggers are listed below, and the first three would change the score rather than only the wording.


Active: the tool is current and recommended.


What Active means here. Active means current and recommended for the work this review describes: research-shaped deliverables, many-item processing, and first-pass builds where you read the result before anything depends on it. It does not mean the product is predictable. The credits a task will consume are not disclosed before it runs, the vendor publishes no independent measurement of reliability, and the two review platforms that carry the most weight for a paid software decision rate it poorly on exactly the issues this review is about. A reader who needs a fixed cost per task, or a measured guarantee, should treat those as unavailable for now and plan around them rather than assume they will improve.


Re-check triggers:


  • The publication of a per-task credit estimate before a task starts, or a hard budget stop. The vendor documents that consumption depends on task complexity and duration and publishes three worked examples, but no surface states what a task will cost before it runs. This is the single change that would move the Time sub-score.

  • An independent, methodologically documented measurement of task reliability. The GAIA figures that circulate for Manus come from commentary rather than from a published run report, and no third party has measured how often a completed task is usable without rework. Until one exists, the Quality sub-score rests on the absence of measurement plus vendor and user reporting.

  • A change to the memory and skills write path. Manus can build a reusable Skill from a completed interaction, and a project workspace carries a standing instruction into every new task created inside it. Any change to whether you approve that write, or whether you can read it, changes the clause 5.2.3-a reasoning and the Skill sub-score.

  • A change to the licensed rights over your content. The terms grant a perpetual and irrevocable licence over Your Content for aggregated use to improve the Services, and the privacy policy separately reserves the right to create aggregated and de-identified data and share it with third parties for lawful business purposes. Any narrowing or widening of either clause changes the adoption calculus in section 7c.

  • A change to the corporate structure or to where tasks are processed. The product's independent status was restored in 2026 after a completed acquisition was unwound, and the privacy policy states that data is transferred internationally and processed by third-party artificial intelligence providers. Both belong in a procurement review for any institution with a data residency rule.

  • A published list of plan prices on the vendor's own pages that a text read can confirm. The vendor's pricing cards render the price in a dynamic element, so the credit allowances can be read from the page and the currency figures can not. Two independent 2026 reviews report the entry price, and they agree with each other and with the vendor's own credit figure. A published rate card would remove the remaining ambiguity.


For detailed explanations of the CI-First evaluation terms used in this review, including the Humics Protection Badge and the AI Imposture Risk levels, see the Glossary at the end of this post.




Back to the TOC

Tool Snapshot


At a Glance Dashboard


Field

Value

Category

Applied AI / Agent Platform

CI-First Benefit Score

5.3 / 10 (CI-First Positive)

Sub-scores

Time 7 / Quantity 6 / Quality 5 / Skill 3

CI-First Profile

Primary: Co-Worker and Assistant (level 2). Secondary: Analyst and Tester (level 4), and narrowly Co-Creator and Thought Partner (level 1)

Collaboration Mode

Centaur. Cyborg is not available on a platform that runs tasks unattended, on a schedule, and concurrently, and that writes durable artefacts you carry forward

Humics Protection

Humics-Risky (-2 / +3): Creativity 0, Critical Thinking -1, Social Authenticity -1

AI Imposture Risk

Medium overall, with Skill Illusion High, Time Illusion Medium, Quantity Illusion Medium

Status

Active

Last tested

2026-09-25

Released

Introduced in early 2025; independent operation restored in 2026 following the unwind of a completed acquisition

Access

Browser workspace at manus.im; desktop, mobile, browser extension, email, Slack and API surfaces; hosted only

Price

Credit based. 4,000, 8,000 and 40,000 credits per month across three paid tiers plus a free tier, with a 17 per cent annual saving. Entry price reported at $20 per month by two independent 2026 reviews

Vendor

Butterfly Effect Pte. Ltd., incorporated in Singapore

Framework version applied

CI-First Evaluation Framework v1.2

Independent measurement

None published in a methodologically documented form at the time of writing


Manus AI (manus.im)


Tagline: "Manus AI is an autonomous general AI agent designed to complete tasks and deliver results." The vendor's own framing is that it "takes action" and operates as "a virtual colleague with its own computer". (manus.im/docs, read 2026-09-25.)


Category: Agent platform. A hosted general-purpose agent that plans and executes multi-step tasks inside a sandboxed virtual machine with internet access, a file system and the ability to install software, and returns finished artefacts rather than answers. The vendor's own term is an autonomous general AI agent.


Primary use cases:


  • Hand over a research brief and receive a structured report or dataset rather than a reading list.

  • Process a long list of similar items with one agent per item, which the vendor calls Wide Research and documents up to 250 items.

  • Produce a first-pass deliverable: a slide deck, a website, a web application, a spreadsheet, a batch of images.

  • Run the same task on a schedule, or run several tasks at once. The paid tiers state 20 concurrent tasks and 20 scheduled tasks.

  • Automate a recurring workflow by saving it as a Skill or as a project workspace with a standing instruction and a reference file set.

  • Act inside your own browser sessions through an extension that uses your existing logins, or inside a vendor-hosted cloud browser that does not.


Pricing summary: A credit system across three paid tiers plus a free tier. The vendor's own plan cards state 4,000 credits per month, 8,000 credits per month and 40,000 credits per month, each including 300 refresh credits every day, 20 concurrent tasks and 20 scheduled tasks, with a 17 per cent saving on annual billing. The free tier is described as a limited monthly credit allowance with access to core capabilities. Two independent reviews published in 2026 put the entry paid tier at $20 per month for the 4,000-credit card and the Team plan at $39 per seat per month with a five-seat minimum and a shared pool of 19,500 credits, and the vendor's own Team page confirms per-seat billing, a shared pool, add-on credits and a 17 per cent annual saving without a readable list figure. Credits are consumed for large language model tokens, virtual machines and third-party API calls, only while a task is running, and the vendor publishes a full refund of consumed credits for tasks that fail for technical reasons on its side.


Official links:



Agent platform fields:


  • Agent architecture: one orchestrating agent per task by default, following a plan, act and observe loop inside a sandboxed machine. A second mode, Wide Research, decomposes a list-shaped task into independent sub-tasks, assigns each to a dedicated agent with its own fresh context, runs them simultaneously and has the main agent assemble the results. The vendor states the main agent collects the completed sub-tasks and builds the requested format.

  • Memory and standing context: three distinct stores. A temporary sandbox that disappears when a task ends. A project workspace that carries a master instruction and an uploaded knowledge base into every new task created inside it, available on all subscription tiers, private by default, with configuration changes not applied retroactively to tasks already created. A persistent cloud machine, created on request, that keeps files and installed tools across tasks. Skills, described below, persist in a library attached to your account.

  • Skills: modular file and folder based instructions that the agent loads on demand, activated with a slash command in a conversation, with four documented routes in: build one with Manus from a successful interaction by instructing it to save the process, upload an archive or folder, take one from the vendor's curated library, or import one from a GitHub repository. The vendor states that community skills can contain code and shell commands and that Manus will audit a skill for you on request.

  • Inputs: a natural-language instruction, optional uploaded files (documents, spreadsheets, images, PDFs, comma-separated values), and, through a project, a standing instruction and a reference file set.

  • Outputs: documents, spreadsheets, slide decks, websites and web applications, images, code, structured tables and datasets, plus executed actions such as browser form filling and scheduled runs.

  • Replay: completed tasks carry a shareable replay link, which is how the vendor demonstrates Wide Research output at scale. It is also the most useful inspection surface the product has, and this review treats it that way.

  • Integrations: Slack, Google Calendar, Gmail, Notion, GitHub, Google Drive, Zapier, MCP connectors and custom MCP servers, a browser extension, and a REST API documented as giving programmatic access to the agent.

  • Platforms: hosted only. There is no self-hosted edition and no open-weights release. Surfaces are the web app, desktop apps, mobile apps, a browser extension, email, Slack, and the API.

  • Model layer: not disclosed per task. The vendor states a design principle of remaining orthogonal to underlying models and relying on context engineering, and the privacy policy names third-party artificial intelligence providers as integrated components. Treat the per-task model as undisclosed.




Back to the TOC

The Problem


Most knowledge work breaks into two halves, and the second half is where the time goes. The first half is deciding what the work should be. The second half is doing it: gathering the material, reading it, structuring it, formatting it, and repeating the same sequence for each item in a list. People who are good at the first half are routinely slow at the second, and the gap widens with volume. A competitor analysis of three companies is an afternoon of thought. A competitor analysis of fifty is a month of assembly, and the thinking is a small fraction of it.


Chat assistants addressed the first half and left the second largely alone. A chat response is a response. It does not open a file, run a script, install a package, visit fifty pages in a browser, build a spreadsheet, or come back in twenty minutes with something you can send. Everything the model produced still had to be moved, formatted and finished by a person, and the moving and finishing was the work that was actually consuming the week.


Two further problems arrive with volume. The first is a context limit that the vendor documents candidly in its own Wide Research page: an assistant working through a list degrades, with detailed analysis for the first few items, shorter descriptions as the context fills, and generic summaries plus increased errors past roughly the tenth, a point the vendor calls the fabrication threshold. The second is that a single long conversation is not a place to run a job. There is no schedule, no concurrency, no file system, and no way to hand off a task and do something else.


The problem an autonomous agent platform addresses is therefore not "how do we get better answers". It is "who finishes the work, and how much of it can run while I am doing something else". The honest question a reader should ask before adopting one is the second half of that sentence: the platform can run unattended, and the work still has to be judged by someone when it comes back.




Back to the TOC

The Outcome


An autonomous agent platform does not remove the work. It moves it from assembling to judging.


What leaves is the assembly: the gathering, the formatting, the repetition, and the mechanical parts of a build. What arrives is reading a finished artefact carefully enough to know whether it is right, deciding what the agent should have done differently, and paying for the run. On a fifty-item research task this is a genuinely favourable trade, because assembly was the cost and judgement is fast when the material is in front of you. On a single high-stakes document it is a much worse trade, because you have replaced the thing you were going to write with the thing you now have to audit, and the audit is harder than the writing would have been.


The vendor's own framing sets an expectation that the second half can be skipped: the product is described as delivering "production-ready results" and as "business AI that works like your best employee". The terms of use say the opposite in the plainest available language. You are responsible for independently reviewing all Output, you are fully responsible for monitoring and approving its use, and the vendor makes no warranty that the output will be accurate or reliable. Both statements are from the same vendor, and a reader adopting the product should act on the second.


So the outcome, stated honestly, is this. You get a second pair of hands that does not get tired, does not need the task explained twice, and will work through a list of two hundred and fifty items at the same depth on the last one as on the first. You get it on a credit meter you cannot read in advance. And you get a finished artefact produced by a system that will not tell you which paragraph it is least sure about, which means the verification burden moves to you and stays there. That last point is what this review returns to throughout, and it is why the Skill sub-score is the lowest of the four.




Back to the TOC

Who Should Use Manus


Learner type

Difficulty

Typical ROI

Career path

Students (Bachelor, Master)

Beginner for running a task, Intermediate for getting a usable artefact

Turning a research brief into a sourced draft, and building a first working web project without a development background. The genuine constraint is cost: at 900 credits for the vendor's own complex example, the free tier covers one such task and not two

Coursework and thesis support, project work, LIPS Collect and CARE Collect phases for organising what comes back

Professionals (career upskilling)

Intermediate

Many-item research and recurring reporting delivered as a finished artefact rather than a reading list. The strongest case is a recurring report that you already know how to write, because you can judge the output and the run can be scheduled

Research, operations and analyst roles, competitive intelligence, programme and product work, and the UDG and UDI adjacent territories

Everyone (lifelong learners)

Intermediate

One place to hand over a multi-step task and get an editable file back. The transferable skill is instruction writing, and it is the same skill every agent platform now needs

LIPS and CARE discipline for capture and review, SL-OS output capture, and a working vocabulary for the agent era


Skill level required: Intermediate. Writing an instruction that states the task, the context, the constraints and the output format is the core skill, and the vendor's documentation says as much in its own Wide Research guidance: be specific about structure, specify the scale upfront, describe the desired output format, include evaluation criteria. A vague instruction produces a vague deliverable at the full credit cost, which is the most expensive way to learn the lesson.


Prerequisites: An account, which is free to create. Two habits make the difference between a useful account and an expensive one. The first is reading the replay and the task log before you trust the artefact. The second is being able to judge the artefact you receive, which is a property of you and not of the product.


Typical time to first result: Minutes. The vendor's own credit examples run 15 minutes for a standard data-analysis and visualisation task, 25 minutes for a standard website build and deployment, and 80 minutes for a complex application build. The work runs unattended, so the time that matters is your writing time plus your reading time rather than your waiting time.


Typical time to competence: Two to four weeks of regular use before instructions reliably produce the artefact you wanted on the first attempt. The specific skill to acquire is stating the output format, because an agent that has not been told the shape of the deliverable will choose one, and it will not be the shape your organisation uses.




Back to the TOC

U365 Institutes Alignment


The table below states the operational competency this tool exercises rather than its subject matter, because a relevance rating without its limit is a label rather than a judgement. No credential or programme claim is asserted anywhere in this review: the rows are alignment readings, and no U365 credential is attached to any of them.


Institute

Relevance

Why

UIT (Technology, AI, Data Science)

Low to Medium

Coursework observation only. The connector layer, the model context protocol surface and the REST API are real integration surfaces a Fellow can work with, and the vendor's own documentation of parallel decomposition is a clear account of why a context window forces an architectural answer. No credential relevance and no credential chain. The competency the review actually identifies, supervising an agent and reviewing the procedure it wrote, has no assessment home, and it would sit in this institute if one were built. The limit that holds the row where it is: the tool produces no system, no assessed code and no data artefact, the integration surfaces are shared with most products in this series, and the vendor publishes no architecture, no weights and no reproducible evaluation surface a Fellow could be assessed against

UIB (Business Management, Entrepreneurship)

Medium

One competency, and it is costable from figures the vendor publishes: the three credit allowances, the worked consumption examples with their credit figures, and the statement that credits are consumed only while a task is processing. A Fellow can run their own brief, count what it consumed, convert the count into a fraction of a monthly allowance and defend a tier, a scope or a route against the number. The review's own credit arithmetic is the worked example. The limit that holds the row at Medium: the vendor publishes no rate card, no per task estimate and no spend cap, so the exercise is cost observation and budgeting rather than rate based appraisal, and the tool teaches no management, finance or entrepreneurship content. The relevance attaches to the cost and budgeting competency only

UIC (Digital Communication, Marketing)

Low to Medium

Coursework observation of a produced communication artefact. The product returns documents, slide decks and landing pages in the institute's own output formats, and the published Content Marketing Specialist programme assesses content strategy and search content writing, so a produced page can be put in front of a Fellow for critique against a brief. No credential relevance and no credential chain. The limit that holds the row: the product produces communication and does not build a communicator. Independent reviewers describe the builds as fast prototypes, and the workflow always needs a human finishing step, so no artefact leaving the tool is assessable as the Fellow's own authored output

UID (Digital Design, UX/UI)

Low

Coursework observation of a produced design artefact. A produced deck or landing page can be put in front of a Fellow for critique against a brief, which is a curriculum exercise in critique. No credential relevance and no credential chain. The decisive limit: the design surface iterates on a composition the agent has already made, so the Fellow reacts to a produced design rather than specifying one, and there is no surface where the human forms the visual decision before the machine produces the artefact. The product makes design material and does not make a designer


Skill level and the honest caveat. The alignment above assumes a reader learning to commission work from an agent rather than to do it by hand, because that is the competency this class of product changes. The skill a reader actually builds is specification and evaluation: stating a task precisely, and judging a finished artefact against the standard you set. That skill transfers to every agent platform. The skill a reader does not build is the subject matter of the artefact itself, which is why this review maps no credential chain for the content disciplines and why the Skill sub-score is 3 rather than 6. No credential is attached to that reading, and a credential chain would have to be verified before one were stated.




Back to the TOC

How Manus Works


Inputs: A natural-language instruction. Optional file uploads: documents, spreadsheets, images, PDFs, comma-separated files. A project workspace, which supplies a master instruction and a file knowledge base to every task created inside it. A connector or MCP server, which supplies a live data source. For the browser extension, your own authenticated sessions and open tabs, which the privacy policy states are transmitted to Manus servers for task processing.


Outputs: A completed task with a deliverable attached. The documented range covers documents, spreadsheets, slide decks, websites and deployed web applications, images, code, structured tables and datasets, and executed actions in a browser. Every task carries a shareable replay, which is both the vendor's demonstration surface and the reader's inspection surface.


The execution loop, in the vendor's own terms: the agent operates in a complete sandbox environment, which the vendor describes as a virtual computer with internet access, a persistent file system, and the ability to install software and create custom tools. It plans, acts, observes the result, and continues. Credits are consumed during that loop against three things the vendor names: large language model tokens for planning, decision making and generation, virtual machines for file operations, browser automation and code execution, and third-party API calls for integrated external data services.


How credits work, and the part that matters to you. The vendor states plainly that the specific credits consumed by a task depend on its complexity and duration, and that credits are only consumed while a task is processing. It publishes three worked examples, which are the most useful pricing document on the site because they are the only figures that let you reason about your own workload:


Example task, as the vendor states it

Complexity

Duration

Credits

Basketball scoring efficiency quadrant chart, data analysis and visualisation

Standard

15 minutes

200

Wedding invitation webpage, design, code development and deployment

Standard

25 minutes

360

Sky events web application with location-based reports, app development, data integration and interactive deployment

Complex

80 minutes

900


Three conclusions follow, and the vendor does not state any of them. First, consumption tracks active minutes closely: the three examples run at roughly 13, 14 and 11 credits per active minute, so the complexity label matters less than how long the machine runs. Second, a 4,000-credit month covers about eleven of the standard website example and about four of the complex application example. Third, and this is the finding, no surface of the product states what a task will cost before it runs. There is no per-task estimate and no spend cap, a gap that the independent reviews and the user complaints in Section 9 both return to. You learn the price by paying it, and the vendor's own refund policy covers only tasks that fail for technical reasons on its side.


Manus: the vendor's three worked credit examples and what the meter does not state before a task runs, illustrating Section 6

How Wide Research works: rather than one agent working a list sequentially, the main agent decomposes the request into independent sub-tasks, assigns each to a dedicated agent with its own context window, runs them in parallel, and assembles the results into the requested format. The vendor documents the mechanism as a direct answer to context degradation, states that quality is uniform at any scale because item 250 gets the same treatment as item 1, and reports testing up to 250 items. It also states where the mode does not fit, which is unusually honest documentation: single deep-dive analysis, tasks with sequential dependencies, real-time interactive work, and any list under ten items.


How skills and standing instructions work, and this is the part of the product a U365 reader should read twice:


  • A Skill is a file and folder based instruction set the agent loads on demand. There are four ways to add one, and the first is the one that matters here: build a skill with Manus itself, which the documentation describes as taking a successful interaction and instructing the agent to save the entire process.

  • A project workspace carries a master instruction and an uploaded knowledge base that apply to every task created inside it. The documentation states that configuration updates do not affect tasks already created, and that members invited to a project share the master instruction and the knowledge base while seeing only the tasks they created themselves.

  • Skills persist in a library and are activated with a slash command, so an artefact created in one session becomes standing operational context in later ones. The vendor's own security note is that community skills can contain code and shell commands and should be verified before use, and that Manus will run the audit for you on request.


Integrations: Slack, Google Calendar, Gmail, Notion, GitHub, Google Drive, Zapier, MCP connectors, custom MCP servers, a browser extension, and a REST API. The API documentation describes sending a task and receiving a complete result, which is the same agent with a programmatic front door.


Platforms and deployment: hosted only. Surfaces are the web application, desktop applications, mobile applications, a browser extension, a dedicated email address for your account, Slack, and the API. There is no self-hosted route and no published offline mode.




Back to the TOC

Getting Started with Manus


Required accounts: A Manus account, created by email or through a Google, Apple or Microsoft sign-in. The free tier is documented as covering core capabilities with a limited monthly credit allowance, and every plan including free carries the 300 daily refresh credits the paid cards state.


Installation: Nothing to install for the core product, which runs in the browser. Optional surfaces are the desktop application, the mobile application, the browser extension, the Slack integration and an API key. Install the browser extension only when you have decided that you want an agent acting inside your logged-in sessions, and read the authorisation step rather than clicking through it.


First-time configuration:


  • Create the account and read the three worked credit examples at manus.im/help/credits before you run anything long. They are the only figures that let you forecast your own consumption.

  • Pick the tier against the credit allowance rather than against the feature list, because every tier carries the same 20 concurrent and 20 scheduled tasks and the free tier states access to core capabilities. The differences that matter are credits per month and the daily refresh allowance.

  • Decide the billing period. Annual billing carries a stated 17 per cent saving, and the monthly credits refresh on your subscription date rather than rolling over.

  • If you will run many-item work, read the Wide Research page before writing your first long instruction, because the mode changes the instruction shape. It wants the scale stated, the columns named, and the output format specified.

  • If you will run recurring work, create a project workspace once and put the standing instruction and the reference files in it, rather than retyping context every week.


First 15 minutes checklist:


  • ☐ Run one small task with a stated output format and read the replay end to end, so you see what the agent actually did rather than only what it produced

  • ☐ Run one list task with between 10 and 20 items and a named column set, and compare the last row's depth against the first

  • ☐ Open the artefact and check three specific facts against their sources, not against your expectations

  • ☐ Note the credits the two tasks consumed and divide by the minutes each ran, so you have your own cost per active minute

  • ☐ Decide, in writing, what class of task you will never hand to it: anything you cannot personally judge is the answer


Result: After 15 minutes you hold one partially verified artefact and one measured cost figure for your own workload. The second item is the one most first-time users never collect, and it is the number that decides whether the subscription is worth renewing.




Back to the TOC

Real Workflows


Workflow 1: The many-item research table, which is the product's clearest strength


The task: Take a list of named entities and produce a structured comparison you can act on, where each row is researched to the same depth.


The instruction:


Context: I am preparing a comparison for [audience and decision]. The list is [n] items and it is in the attached file. The fields I need are [name four to six columns precisely]. I already know [state what you know], so do not spend effort re-establishing it. Profile: Act as an Analyst and Tester (level 4). I will interpret and decide. You gather, structure and flag uncertainty. Task: Research each item in the attached list and produce one table with one row per item and the columns named above. For every cell, cite the source you used. Where you cannot establish a field from a source, write "not established" rather than inferring it. Constraints: Use the web sources you can reach, and prefer primary sources such as the company's own site or a filing over a directory listing. Do not fill a cell by analogy to a similar item. Do not merge two items into one row. State the date you read each source. Output format: A single table with the named columns, a source column, and a date-read column. Then a short list of the items where you could not establish two or more fields, with the reason.


Why the constraint line does the work: an agent left to its own judgement will fill a gap with a plausible value, and a plausible value in a comparison table is indistinguishable from a researched one. "Not established" is the single most valuable output the agent can produce, because it tells you where your remaining work is.


Verification checklist. Applied to the many-item research table.


  • ☐ Multi-Model Check: run the same list through a second agent or assistant on a sample of five rows and compare the values that differ

  • ☐ External Source: open the cited source for at least three rows, chosen at random rather than chosen as the ones that look least reliable

  • ☐ Human Review: have the person who will act on the comparison read the three rows closest to the decision

  • ☐ CI-First Test: can you explain and defend every cell in the row that drives your decision without the tool? If not, that row is not ready


Workflow 2: The recurring report you already know how to write


The task: Turn a report you produce on a known cadence into a scheduled run with a standing instruction, so the assembly disappears and the judgement stays with you.


The instruction, set once in a project workspace:


Context: This project produces the [name] report [weekly or monthly]. The audience is [name the reader and what they decide]. The sections are [list them]. Our house format is [state the heading structure, the length and the tone]. The reference material is the knowledge base attached to this project, plus [name the live sources]. Profile: Act as a Co-Worker and Assistant (level 2). You assemble and draft. I own the analysis and every sentence that carries a recommendation. Task: Produce the current edition following the house format. Where a number has moved materially since the last edition, say so and state the previous number. Where you could not reach a source, say which one and what you used instead. Constraints: Do not write recommendations. Do not smooth over a gap in the data. Do not carry forward a figure from the previous edition without re-reading the source. Flag anything you infer rather than read. Output format: The documented section structure, then a short "what I could not establish" list, then a one-line note naming every source that failed to load.


The honest caution: the value here is real and it is also where the credit model bites, because a recurring report is a recurring cost at a price you do not know in advance. Run it manually twice and record the consumption before you schedule it.


Verification checklist. Applied to the recurring report.


  • ☐ Multi-Model Check: not required for format, required for any figure you will publish

  • ☐ External Source: verify every number that changed since the last edition against the underlying source yourself

  • ☐ Human Review: you, reading the whole thing, before it goes anywhere. This step does not automate

  • ☐ CI-First Test: could you write this report from scratch if the subscription were cancelled tomorrow? If not, the tool has taken the skill rather than the labour


Workflow 3: The first-pass build, where the value and the risk sit closest together


The task: A landing page, a small web tool, or a deck you intend to finish yourself.


The instruction:


Context: I need a [artefact] for [purpose and audience]. It must feel like [describe the reference or the existing brand], and the constraints are [hosting, size, data, accessibility, anything the organisation imposes]. Profile: Act as a Co-Worker and Assistant (level 2). I own the judgement about whether it is good. You produce the first pass. Task: Build it and deploy a preview I can open. Then tell me the three decisions you made that I did not specify, and why. Constraints: Do not use placeholder text or lorem ipsum. Do not invent a statistic, a customer name or a testimonial. Do not use an image you cannot tell me the licence of. If a requirement is ambiguous, state the reading you chose and continue rather than stopping. Output format: The live preview link, a list of the three unspecified decisions with your reasoning, and a list of the requirements you could not satisfy.


Why the third output line matters: the three unspecified decisions are where a build stops being yours. Reading that list is the difference between commissioning work and discovering it.


Verification checklist. Applied to the first-pass build.


  • ☐ Multi-Model Check: not applicable to a build; applicable to any copy inside it

  • ☐ External Source: open the deployed page on a real device, click every interactive element, and check every claim against a source you can name

  • ☐ Human Review: a colleague who did not commission it, looking for the things you are now blind to

  • ☐ CI-First Test: can you maintain this artefact when the agent is not available? If not, say so in the handover record rather than assuming you can




Back to the TOC

Strengths, Limits, and AI Imposture Risk


What Manus does better than the alternatives


The many-item case is genuinely differentiated, and the mechanism is architectural rather than promotional. Most assistants degrade as a list grows, for the reason the vendor documents on its own Wide Research page: one context window has to hold every item processed so far. Manus answers that with one dedicated agent and one fresh context per item, running in parallel, with the main agent assembling the result. The vendor reports testing to 250 items. Whether it performs at that scale is a separate question this review treats in the Quality dimension, and the design is the correct answer to the stated problem regardless of the performance question.


The breadth of finished artefacts is unusual. A single platform that returns a report, a spreadsheet, a deck, a deployed website, a working web application, a batch of images and a code repository is not a thin wrapper around a text model. The website builder is documented down to hosting, custom domains, third-party payments, analytics and code export, which is a product rather than a demo.


The documentation is candid in places where vendors are usually not. The Wide Research page states where the mode does not fit, including any list under ten items and anything with sequential dependencies. The credits page states the three cost drivers and admits consumption is a function of complexity and duration. The terms of use state the limits of the technology more bluntly than most terms documents do, including that outputs may contain errors, that AI cannot understand or express emotions as humans do, and that you are fully responsible for monitoring and approving the use of output.


Unattended execution and the schedule are real, and they are what makes the Time case work. The paid tiers carry 20 concurrent tasks and 20 scheduled tasks, so the platform is a place to leave work rather than a window to sit in front of.


The limits


You cannot know what a task will cost before you run it. This is the single most consequential limit for a reader with a budget. The vendor publishes three worked examples and no estimate, no cap and no pre-run quote. Independent reviews, user reviews and the vendor's own consumption documentation all converge on the point that the cost arrives after the fact. A 4,000-credit month covers roughly eleven runs of the vendor's own standard website example and roughly four of its complex application example, and a run you abandon halfway still consumed its credits.


Independent verification of reliability does not exist. The GAIA benchmark result that circulates for Manus, including the widely repeated figure of 86.5 per cent at level one, is quoted in commentary and academic overviews rather than published as a vendor run report with its configuration and date, and no third party has measured how often a completed agent task is usable without rework. This is not a claim that the product is unreliable. It is a statement that nobody has published the measurement either way, and this review scores Quality down for that absence and says so at the number itself.


The product is not a document editor and the outputs are not finished. On a long report, revising one section means re-running or rebuilding it, because the artefact is the output of a run rather than a document you are working in. Independent reviewers are consistent on this point and so is the vendor's own positioning: the value is the first pass.


The verification burden lands on the person least able to carry it. An agent that can produce a financial model, a legal-flavoured summary, a piece of code or a compliance table for a user who cannot evaluate any of the four is not offering a skill. It is offering the appearance of one, and the product's design makes that appearance unusually convincing because the output arrives finished, formatted and deployed.


Skills and project instructions become durable standing context, and this is the mechanism behind the High Skill Illusion rating. The documentation states two routes by which the platform places instructions into future runs. A project workspace applies a master instruction and a knowledge base to every task created inside it, and it is available on every tier including free. A Skill can be built by Manus from a successful interaction by instructing it to save the process, and it then persists in a library and activates by slash command. Both are durable. Both are reused later. And the first of them, the project master instruction, is the one a careful user writes themselves, while the Skill path is the one where the agent is doing the authoring on your behalf.


Concentration and provenance. The contracting entity is Butterfly Effect Pte. Ltd., incorporated in Singapore, and the privacy policy states that personal information is transferred internationally and that third-party artificial intelligence providers are integrated into the service. In 2026 the product's parent returned to independent operation after an acquisition that had already completed was ordered unwound by China's National Development and Reform Commission, on grounds resting on the Chinese origin of the technology rather than the location of the company. That is published reporting rather than a finding of this review, and it is a fact a university reader should have in front of them when deciding where a research corpus or a student project is processed. Section 7c sets out the contract terms that govern the data position.


CI-First Profile Classification


Primary profile: Co-Worker and Assistant (level 2). The product's value proposition is stated by its own vendor as taking action and delivering results: it plans, executes and produces deliverable work products while you supervise and judge. That is delegation with review, which the framework places at level 2. It is not level 1, because co-creation and thought partnership describe a human and a model building on each other's thinking toward a direction neither had fixed, and Manus is designed to be pointed at a task you have already defined and to return a finished artefact rather than to iterate a direction with you.


Secondary profile: Analyst and Tester (level 4), and this is a real second mode rather than a formality. The tool's documented value includes finding what you missed, through the breadth of source reading it can do in parallel and through the explicit "what I could not establish" reporting that the workflow instructions in Section 6 ask it to produce. The same profile was assigned as a secondary for Klarent, and the test is identical: does the tool's value include analysis the user would otherwise have to perform, and here it does.


Narrow third profile: Co-Creator and Thought Partner (level 1), narrowly. Design View provides a real iterative surface where you give feedback on a visual and it changes, and Wide Research can be used to survey a question rather than to fill a table. The rating is narrow because the product's dominant mode is commission-and-deliver rather than think-alongside.


What does not fit. Coach and Tutor (level 3) does not apply. Nothing in the product is designed to teach you the subject matter of what it produces, and its own documentation offers no learning mode. The one place a reader learns something is by reading the replay and asking why the agent chose what it chose, which is a discipline you impose rather than a feature the vendor built. Challenger and Devil's Advocate (level 5) does not apply either: the product will report what it could not establish, and it will not argue against your framing of the task.


Collaboration Mode


Recommended mode: Centaur.


Alternative mode: None recommended. Cyborg is not available on this product.


Mode rationale: Two independent grounds, and both belong on the record. The first is the framework's own rule at Section 7.2: Centaur is assigned when the Imposture Risk is Medium or High, and this product is Medium overall with Skill Illusion High. The second is specific to an agent platform. Cyborg requires a fast iteration loop with a stopping criterion the human applies inside it. Manus runs unattended, on a schedule, up to twenty tasks at once, and it writes durable artefacts into your library and your projects. There is no loop to stop inside in the sense Cyborg needs, and the writing continues after you have stopped watching. The boundary that makes Centaur real here is procedural rather than interactive: define the task boundary and the output format before the run, read the replay when the task is consequential, keep the master instruction of every project in your own hands rather than delegating the writing of it, and review what a Skill contains before it enters your library. You own what counts as done and what the deliverable must contain. The platform owns the execution.


CI-First Benefit Score


Dimension

Score (0-10)

Rationale

Time

7

Strong savings on the task types this product fits. The mechanism is structural: you write one instruction and the platform runs for minutes or an hour without you, which is work that previously consumed the whole of your attention even though most of it was mechanical. The vendor's own worked examples put a standard website build and deployment at 25 minutes and a complex application build at 80 minutes with no human time inside either. Held below 8 for two honest reasons. First, the overhead is not zero: writing an instruction precise enough to produce a usable artefact takes practice, the first attempts take longer than doing the work, and two to four weeks of use is a realistic ramp. Second, and unlike a task you can time yourself, the cost is metered and unstated in advance, so a run that goes wrong has cost you something you cannot recover. This is a strong Time score with a spend risk attached to it, and both halves belong in the number

Quantity

6

Moderate to strong increase, and the strongest single case for the product. The parallel decomposition model is a real multiplier rather than a productivity claim: the vendor documents testing to 250 items with one dedicated agent per item and uniform depth at scale, and the same platform can run 20 tasks concurrently and 20 on a schedule. A reader who previously assembled a twenty-entity comparison by hand, or produced one report a month, genuinely produces more usable output. Held below 7 because the scored quantity is verified and usable volume, and the verification step scales with the volume: a fifty-row table requires fifty rows of checking from a person, and the honest user checks far fewer than fifty. The quantity is real and the capacity to judge it has not grown with it, which is the Quantity Illusion in its ordinary form

Quality

5

This is the dimension the review turns on, and the score is set for the absence of measurement, stated plainly rather than implied. No independently measured, methodologically documented assessment of Manus task reliability exists: the GAIA benchmark figures in circulation are quoted in commentary and academic overviews rather than published as a dated vendor run report with its configuration, and no third party has measured how often a completed task is usable without rework. Against that, the product does carry genuine quality mechanisms: parallel per-item context removes the degradation mechanism the vendor documents, every task carries a replay so the process is inspectable rather than only the result, the terms place responsibility for reviewing output squarely on the user, and independent reviewers report satisfaction among users with defined, research-shaped tasks and dissatisfaction among users expecting polished production output. The number is 5, which is the Moderate band on the framework's 0 to 10 scale and the top half of the range that maps to CI-First Positive; 5 is what an unmeasured but genuinely capable tool scores, and it is not a rejection. A quality claim without a published measurement is not a measured quality benefit, and raising this above 6 would assert a reliability figure nobody has produced

Skill

3

Marginal benefit, scored conservatively as the framework directs, and the score is a judgement under the framework's benefit rubric rather than a clause floor. Clause 5.2.3-a sets the floor on the Skill Illusion rating, which is High, and it does not set this score. There is a real transferable capability, and it is instruction writing: stating context, profile, task, constraints and output format is a skill that transfers to every agent platform, and it is the skill a reader builds fastest here. Beyond that the product is skill substitution by design. It produces the deliverable and teaches you nothing about its subject, and the person best positioned to commission a financial model, a compliance table or a code module from it is a person who can already judge those things, while the person who cannot is the one most likely to trust the result. Clause 5.2.3-a then applies its floor: the platform both builds reusable Skills from completed interactions on the user's behalf and carries project master instructions forward into every future task, so the user holds documented capability they did not author, and the Skill Illusion is therefore High. That finding is about the illusion the tool creates in the user; the benefit score of 3 below stands on its own reasons, which are the substitution of output for capability and the low transferable yield beyond instruction writing. Skill 3 is the Marginal band. The number is about what the tool builds in you, not about what the tool produces for you


CI-First Benefit Score: (7 + 6 + 5 + 3) / 4 = 5.25, rounded to 5.3 / 10 (CI-First Positive)


Why this score is not higher, and why it is not lower


5.3 is CI-First Positive. The band label matters and it should be read with the number: this is a recommendation for the work it fits, not a caution. A person or team commissioning research-shaped and many-item deliverables on a recurring basis has a genuine and measurable net benefit available here, and two of the four dimensions are strong.


The score is not higher for one reason that appears twice. The framework measures benefit net of overhead, and the overhead here is verification plus an unreadable meter. Nothing independent measures whether tasks complete correctly. The cost of a run is disclosed after it happens. A tool that produces finished artefacts you cannot always judge, at a price you learn afterwards, cannot be scored in the Strong band on the framework's rules however impressive the demonstration is, and the demonstration is impressive. It is also worth saying that the two dimensions holding the total down, Quality at 5 and Skill at 3, are the two where a reader's own practice changes the outcome, which is the point of the framework rather than a defect in the tool.


The score is not lower because the core capability is real and the architecture is the reason. Parallel per-item decomposition solves a problem every other assistant still has. The artefact range is genuine rather than a wrapper. Unattended and scheduled execution is real. The vendor documents its own mode's limits and its own cost drivers, and the terms state the technology's limits in language more candid than most vendors use. Time 7 and Quantity 6 are honest middle-to-high scores for a product that genuinely removes assembly work, and the two capped dimensions are capped on measurement grounds and on what the tool builds in the user, not on any observed failure of the product to do what it says.


Humics Protection Badge


Dimension

Rating

Rationale

Creativity

Neutral (0)

The product executes a task you defined. Where it does originate, it originates inside your brief rather than a direction of its own: Design View iterates on your feedback, and Wide Research surveys a field you chose. It neither replaces your ideation in the common case nor trains your creative judgement, so the framework's answer is neutral rather than erosion. There is a genuine risk to name anyway, and it belongs in the Limits rather than in this rating: a person who commissions every first draft stops writing first drafts, and the ability to originate is exercised by originating

Critical Thinking

Erodes (-1)

Four mechanisms, all documented. First, the output arrives finished, formatted and often deployed, which is the presentation most likely to be accepted without inspection. Second, the vendor's own terms place the entire review burden on the user and state in the same document that you are fully responsible for monitoring and approving the use of output, which is a warning placed where few users read it. Third, the cost model rewards trusting the run: since the credits are already spent when the task completes, there is a real temptation to accept the result rather than pay to redo it. Fourth, the two review platforms with the most weight for a paid decision report that the dominant complaints are billing and credit opacity, which tells you where users' attention goes rather than where it should. The mitigation is available and cheap, and that is why this is -1 rather than a more severe reading: read the replay, check three facts against sources that are not the tool, and write your own project master instruction rather than letting one be written for you

Social Authenticity

Erodes (-1)

This is the most consequential change from the null this series has recorded for several tools, and it rests on a vendor feature rather than an inference. Mail Manus is documented as an email-based task interface for the agent, and the browser extension is documented as acting on your behalf inside web pages you are logged into, including form filling, and the privacy policy states that browser actions including clicking, scrolling, form filling and navigation are performed on your behalf and that session context leverages your existing browser sessions and login states. Framework clause 4.2-a is explicit that agent-mediated conversation is not erosion by itself, and that erosion requires either agent-authored text presented as the person's own voice in a human-facing channel, or the substitution of agent interaction for human contact. Manus meets the first condition in the general case: text and content produced by an agent that acts through your authenticated identity, sent to humans who have no way of knowing it was not you, is the clause's first condition almost exactly. The rating is -1 rather than a more severe reading because the product does not make this the default, the vendor documents the authorisation step, and a user who keeps the agent inside its own workspace and away from their mail and their logged-in sessions does not touch this dimension at all. The clause note records the mechanism rather than the conclusion


Humics Protection Score: 0 + (-1) + (-1) = -2 / +3 Badge: Humics-Risky


This is the second Humics-Risky badge in the series, after Rabbit OS3, and it should be read with its reason attached rather than as a verdict on the product. The erosion is in two places and neither is a defect of the engineering. Critical thinking erodes because a finished artefact is the easiest thing in the world to accept, and the meter makes redoing it expensive. Social authenticity erodes because the platform is designed to act under your identity, in your mail and in your logged-in sessions, and the people receiving that output cannot tell. Every part of both has a control attached, and the controls are set out in the guidance below. Humics-Risky describes a tool that erodes two capabilities if you let it, and the whole point of the Co-Intelligence framework is that the choice stays with you.


Superhuman Usage Guidance


When to invite this tool:


  • Many-item research where the deliverable is a table, a dataset or a comparison, and where you can verify the fields that matter. This is the product's clearest strength and the case where its architecture does something no chat assistant does.

  • A recurring report or analysis you already know how to produce and can personally judge. Scheduling it removes the assembly, and your judgement stays the constraint rather than disappearing.

  • A first-pass build you intend to finish: a landing page, a prototype, a deck, a small internal tool. Treat the output as a draft you commissioned rather than a product you bought.

  • The parts of a project that are gathering and formatting rather than deciding, in a task whose output you will read closely anyway.

  • A task where the honest answer to "could I do this without the tool" is yes, and the question is only whether you should spend the afternoon on it.


When to keep this tool out:


  • Anything you cannot personally evaluate at the standard the output will be used at. This is the Skill Illusion case and it is the single most important line in this review. A financial model, a compliance table, a legal summary or a security assessment produced for someone who cannot audit it is a liability with a nice format.

  • Any task where a wrong answer is worse than no answer and the verification cost exceeds the doing cost. The framework's Executive Safeguard applies: assume the worst available output and decide whether your check would catch it.

  • Mail and messaging sent under your identity. The product can act inside your authenticated sessions, and this review's Social Authenticity rating is the reason not to point it at accounts where human recipients would reasonably assume the words are yours.

  • Anything you would send to a human without reading every word, on a task whose subject matter you do not know.

  • Data you are not authorised to send to a hosted service in Singapore with international transfer and third-party artificial intelligence providers involved. Read Section 7c before you put a research corpus, a student record or a client document into it.

  • A workflow whose cost must be predictable. There is no pre-run estimate and no spend cap, and a recurring job on a metered, unstated cost is a budget risk rather than a productivity gain.

  • A task where the point of doing it is the skill you build by doing it. Delegating the thing you are trying to learn is the clearest way to arrive at the end of a course having learned nothing.


U365 method integration:


  • LIPS + CARE: the natural division is between Collect and Review. Let the platform gather and assemble, and keep the Review phase as your own work, with what you decided about the artefact recorded in your Digital Second Brain rather than only in the tool's task list. Your research decisions are an institutional asset and they belong where you own them.

  • ULM + EVA: the honest reading is Career and Quality of Life, in that time released from assembly is time returned to work you would rather do. Character is touched in one specific way: refusing to ship an artefact you cannot defend is a discipline, and the product makes the opposite choice easy by making the artefact look complete.

  • UP-Context: this tool needs the full method, with the constraints line doing the most work. An under-specified instruction produces a confident, well-formatted, wrong deliverable at the full credit cost, and the vendor's own guidance says the same thing in its own words when it tells you to state scale, structure and output format. The prompt pack in Section 11 uses the method's order.

  • SL-OS: relevant at the output layer. Reports, decks, tables and builds are files, and files belong in the same places your other work lives rather than only inside the tool's workspace. The scheduled-task surface is the closest the product comes to a routine, and it fits a review cadence rather than a daily one.

  • UNOP: no direct fit as a product. The indirect use is the replay: reading a completed task's replay and asking why the agent chose each step is a deliberate-recall exercise on how a task decomposes, which is a genuine learning material the product hands you without charging extra.


Over-delegation warning. Two failure modes, and the second is specific to this product.


The general one is that the product's ease is the risk. Describing a task costs nothing, the run happens without you, and the result arrives as a file. The CI-First formula applies as it always does: if your Human Intelligence input falls while the Artificial Intelligence term rises, the product falls. With an agent platform the fall is quiet, because nothing breaks, the artefacts keep arriving, and the only symptom is that you have stopped being able to tell a good one from a bad one.


The specific one is the standing instruction, and it is the most consequential finding in this review. A project workspace carries a master instruction into every task created inside it, and a Skill built from a completed interaction persists in your library and loads on demand. Both are durable. Both are reused. The first is usually written by you, and the second can be written by the agent. That is the exact mechanism framework clause 5.2.3-a exists to catch: at the end of it you hold a documented capability you did not author, that outlives the task that produced it, and that you will act on in a later session. The honest practice is to read every Skill before it enters your library, to write the master instruction of every project yourself, and to keep a habit of opening a project's standing instruction and asking whether you still agree with it. A user who treats the library as an accumulating asset they own will do well with this product. A user who treats it as a place where saved workflows appear has delegated the design of their own working method to a system that was never asked to have one.




Back to the TOC

Section 7c: Terms, training statements and permissions


This section exists because an agent platform is a contract before it is a product, and because the gap between what a vendor's marketing pages say and what its terms say is where a reader's real risk sits. This is a contract terms application of the section, on the same footing as the Klarent review, rather than a supplier-conduct application. Everything below is quoted from the vendor's own published documents, read on 2026-09-25. Allegation is kept separate from finding, and at the end of the section it is stated plainly that no sub-score changed and why.


Who you are contracting with, and which document governs. The Terms of Use state the counterparty: "This Agreement forms a legally binding contract between you ("User", "you", "your") and Butterfly Effect Pte. Ltd. ("Company", "we", "us", "our")." The Terms are dated 28 November 2025. A second document governs the Team plan, the API and other business services, and the Terms state the hierarchy directly: "Our Master Services Agreement governs the use of Manus Team Plan, our APIs, and our other services for businesses and developers. If you use these commercial services, your use will be subject to the separate Manus Master Services Agreement and the corresponding Order Form." A reader evaluating Manus for an institution should note which document applies to them, because this review read the Terms of Use and the privacy policy, and the Master Services Agreement governs the commercial route. The Master Services Agreement page returned a title and no readable body on the date read, so nothing in this review rests on its content.


The licensed rights over your content, which is the operative finding of this section. The Terms grant a broad licence over what you put in, and the scope is worth reading slowly:


"You grant us and our affiliates, successors, and assigns a non-exclusive, worldwide, royalty-free, fully paid-up, transferable, sublicensable (through multiple tiers of direct or indirect authorization) right to: (a) during your use of the Services, allow us to copy, display, upload, perform, distribute, store, modify, and otherwise use Your Content to provide and operate the Services and monitor your compliance with these Terms; and (b) a perpetual and irrevocable license to use Your Content in an aggregated manner to improve the Services (e.g., hosting the website, generating AI content at the User's request) and create Usage Data."


Two parts of that sentence deserve to be separated. Sub-clause (a) is ordinary service operation and is time-limited to your use of the Services. Sub-clause (b) is not: it is perpetual, irrevocable, and covers Your Content used "in an aggregated manner to improve the Services" and to create Usage Data. The same passage then limits the reach of sub-clause (b) in a way a reader should hold onto: "For the avoidance of doubt, our use under the foregoing license does not imply our endorsement or ownership of Your Content in any event." So the vendor is not claiming to own your content, and it is claiming a perpetual right to use it in aggregated form. The vendor also states the ownership position plainly elsewhere in the same document: "Please note, we do not own any of your Input or Output." What it does retain is "the Usage Data (as defined below), the Services (including the skills, expertise, and methods used to provide the service), and any improvements, enhancements, or modifications thereof".


The privacy policy's parallel statement, which is a different instrument saying a compatible thing. The policy sets out a purpose of processing under the heading "To create aggregated, de-identified and/or anonymized data", and states: "We may create aggregated, de-identified and/or anonymized data from your personal information and that of other individuals whose personal information we collect... We may use this aggregated, de-identified and/or anonymized data and share it with third parties for our lawful business purposes, including to analyze and improve the Service and promote our business." Its lawful-basis table lists "Research and development / To create aggregated, de-identified and/or anonymized data" with the basis as legitimate interest. The two documents are consistent with each other, and they should be read together rather than as a contradiction: the terms describe the licence you grant, and the policy describes the processing the company performs. Neither states that your inputs are used to train a foundation model, and neither rules it out in those words. The honest reading is that the vendor reserves a perpetual aggregated-use right and a third-party sharing right over de-identified data, and does not publish a statement either way about foundation-model training on customer inputs.


The responsibility allocation, which is unambiguous and belongs in front of any institutional adopter. Under a heading of its own, "1.2 User Responsibilities", the Terms state: "You are responsible for independently reviewing all Output (as defined below). You should exercise personal judgment before relying on Output. You are fully responsible for monitoring and approving the use of Output. You assume responsibility for any decisions, actions, or omissions based on Output." The same document states that AI systems "are based on probabilistic models, which may result in misunderstandings or errors", that the Company "is not responsible for any misunderstandings or inaccuracies caused by AI", and, in the disclaimer, that the vendor makes no warranty that "THE OUTPUT, ADVICE, RESULTS, OR INFORMATION, WHETHER ORAL OR WRITTEN, OBTAINED FROM USE OF THE SERVICES WILL BE ACCURATE OR RELIABLE". It also states that output "may contain a 'Made with Manus' watermark or other forms of identification, which are inherent components of the system and cannot be removed at this time", which is a provenance disclosure rather than a defect.


Permissions the product asks for, stated by the vendor. Three surfaces matter. The first is the browser extension and its associated connector. The privacy policy describes what that grants: page content that Manus reads, extracts and processes from pages you authorise, "including content from premium or authenticated services you are logged into"; browser actions such as clicking, scrolling, form filling and navigation "performed on your behalf"; and "session context leveraging your existing browser sessions and login states", with the data transmitted to Manus servers for task processing and no storage of login credentials. The same passage states the revocation path: access can be withdrawn at any time by disabling the connector or uninstalling the extension. The second is the sandbox, which the policy describes as isolated virtual machines, and which the vendor states collects files you upload or Manus creates, shell commands and their outputs, code generated during the run, and task execution logs for debugging and security purposes. A virtual machine that can install software and run shell commands is the reason the vendor's own Skills documentation warns that community skills "can contain code and shell commands" and should be verified before use. The third is the data position: the policy states that personal information is transferred internationally, that the entity is incorporated in Singapore, and that "Third-Party AI Providers" are integrated into the service and may infer information about you and share information with the vendor as part of processing your inputs and prompts.


Corporate history, stated as reporting rather than as a finding. The vendor's own notice, "Manus Resumes Independent Operations", states that the company has resumed independent operations under its founding team and that for some users the transition "required backing up and restoring data, as well as navigating a temporary interruption to access". Independent press reporting in 2026 states that an acquisition of the company by Meta, announced in December 2025 and described as already completed, was ordered unwound by China's National Development and Reform Commission under its foreign-investment security review, on grounds resting on the Chinese origin of the technology rather than the location of the company, and that the unwind required data isolation and deletion. That reporting is cited here as published reporting, not as a verified fact of this review, and it is material to a reader because it bears on where a data-processing relationship sits and how stable it will be. It is not a finding about the product's behaviour.


What this section does not do, and no sub-score changed. Nothing in this section moved a score. Time, Quantity, Quality and Skill stand at 7, 6, 5 and 3, and the Humics ratings stand at 0, -1 and -1. The reason is the framework's own construction. The CI-First Benefit Score measures benefit to the human net of overhead, and the Humics rating measures the effect of use on three human capabilities. Supplier contract terms are not dimensions of either, which the framework's own review of supplier conduct for MiMo-V2.6-Flash recorded as a deliberate blind spot rather than an oversight. The clause reasoning above did change a rating, and it changed it on a mechanism, not on a term: clause 4.2-a applies because the product is documented as acting under the user's identity in the user's own authenticated sessions, which is a product capability and would be the same finding if the contract said nothing about it. The terms matter here for a different and more practical reason. They are where a U365 reader learns that the review obligation the terms of use and the privacy policy place on a university adopter is larger than the product's marketing language implies, and they are where a reader with a data residency rule learns that the processing is international and involves third-party model providers. This section does not constitute legal advice, it does not assess the enforceability or the fairness of any clause, and it does not allege that the vendor does anything beyond what it states in its own documents.




Back to the TOC

U365 Co-Intelligence Rating


CI-First Evaluation Summary


Field

Value

Tool

Manus AI (manus.im)

Vendor

Butterfly Effect Pte. Ltd., incorporated in Singapore

Category

Applied AI / Agent Platform

Version reviewed

Manus, as documented at manus.im in September 2026

Framework

CI-First Evaluation Framework v1.2

CI-First Benefit Score

5.3 / 10

Band

CI-First Positive (4.1 to 6.0)

Time

7 / 10 (Strong)

Quantity

6 / 10 (Moderate)

Quality

5 / 10 (Moderate, capped by the absence of independent measurement)

Skill

3 / 10 (Marginal, and the Skill Illusion rating is High under clause 5.2.3-a)

CI-First Profile

Primary Co-Worker and Assistant (level 2); secondary Analyst and Tester (level 4); narrowly Co-Creator and Thought Partner (level 1)

Collaboration Mode

Centaur

Humics Protection

Humics-Risky at -2 / +3: Creativity 0, Critical Thinking -1, Social Authenticity -1

AI Imposture Risk

Medium overall: Time Illusion Medium, Quantity Illusion Medium, Skill Illusion High

Status

Active

Last tested

2026-09-25


Framework v1.2 clause note


Clause 5.2.3-a, agent-authored procedural memory: applies. Skill Illusion is High. The clause states that agent-authored procedural memory is a Skill Illusion vector in its own right, that the rating is no lower than Medium for any tool that writes procedural memory on the user's behalf even where a write-approval gate exists, and that the rating is High where the agent can revise that memory during use without a per-write human decision. Two documented mechanisms put Manus at the High threshold. The first is the Skills system: the vendor documents a path in which you instruct the agent to save the entire process from a successful interaction, which creates a durable file-based Skill that persists in a library and loads on demand in later sessions. That is memory written by the agent rather than by the user. The second is the project workspace, which applies a master instruction and a knowledge base to every task created inside it, is available on all tiers including free, and is described by the vendor as removing the need for repetitive setup. The clause's High condition is met by the first mechanism specifically: a Skill created from a completed interaction is not approved per write in the sense the clause requires, because the write and the approval are the same instruction, and the artefact then governs later runs. The mitigation the clause contemplates is that the user reads what was written, and the honest position is that most users will not. This is the strongest application of the clause in the series so far, because the durable artefact is more than a note: it is a procedure, and a procedure that is wrong will be executed rather than merely read.


Manus: the two durable artefacts, a project master instruction and a Skill built from a completed interaction, and how each governs later runs, illustrating the Framework v1.2 clause note on clause 5.2.3-a

Clause 4.2-a, agent-mediated conversation: applies, on its first condition. Social Authenticity erodes. The clause states that agent-mediated conversation is not erosion by itself, that it measures the human's own communicative capability rather than the composition of the channel, and that erosion requires either agent-authored text presented as the person's own voice in a human-facing channel, or the substitution of agent interaction for human contact. Manus meets the first condition. Mail Manus is documented as an email-based task interface, and the browser extension is documented as acting inside pages the user is logged into, with the privacy policy stating that browser actions including clicking, scrolling, form filling and navigation are performed on the user's behalf, that session context leverages existing browser sessions and login states, and that this data is transmitted to Manus servers for task processing. An agent that composes and sends under the user's own authenticated identity, to recipients with no way of distinguishing it from the user, is agent-authored text presented as the person's own voice in a human-facing channel. The clause's second condition is not met on the evidence read, because the product does not substitute agent interaction for human contact; it acts in the user's existing human channels. That distinction is why the rating is -1 rather than more severe. The null this series recorded for Klarent and Rabbit OS3 does not hold here: those products produced platform-attributed reports, and this one is designed to act under the user's identity.


Clause 7.5, team-level rooms: returns a null. The clause covers a room shared by several named agents and the human, requires a profile attributed to each agent individually, requires Centaur as the default, and makes Cyborg unavailable for any room with more than one agent. Manus does run several agents in one execution, and Wide Research deploys what the vendor describes as hundreds of independent agents on a single task. That is the distinction the clause turns on, and the task states it: a distinction between a room where several named agents and a human share a channel, and one orchestrator running subagents inside a single execution. Wide Research is the second case. The sub-agents are spawned by the main agent, they are not persistent named participants, they do not address each other or the human, the human has no per-agent channel, and the results are returned by the main agent as one assembly. Manus Collab, the closest surface to a shared room, is documented as real-time collaboration between humans on a task rather than a channel of named agents. So clause 7.5 returns a null with its mechanism stated: multiple agents inside one execution under one orchestrator, no shared channel with named participants. The clause would apply if the product exposed named persistent agents that a human and other agents could each address in a shared space, and that is the change to re-check for.


Collaboration Mode statement


The required mode is Centaur. The framework's rule at Section 7.2 assigns Centaur whenever the Imposture Risk is Medium or High, and Manus is Medium overall with Skill Illusion High, so Centaur is required rather than preferred. The product-specific ground reinforces it: Cyborg depends on a stopping criterion the human applies inside a fast iteration loop, and an agent platform that runs unattended, on a schedule, up to twenty tasks at once, and writes durable artefacts into your library after you have stopped watching has no such loop. Cyborg is not available on this product.




Back to the TOC

What Users Say


Manus is unusual in this series in that there is a large independent review corpus, and it is unusually negative on the two issues this review cares about most. That is the finding, and this section reports it with the platforms' own numbers and counts rather than as a sentiment.


Aggregate Rating Table


Platform

Rating

Number of reviews

Link

Google Play

4.6 / 5

More than 450,000 ratings

App Store

4.74 / 5, by the reading of one aggregator

Approximately 34,900

Trustpilot

1.2 / 5, rated "Bad"

94 reviews on one regional reading of the page, with one independent review citing approximately 160

G2

2.7 / 5

7

Product Hunt

4.4 / 5, as quoted in an independent 2026 review

9, as quoted in the same review

No listing located at the address checked on the date read

Aggregator reading across three app and web sources

4.60 / 5

443,400 across the three sources, of which approximately 408,500 is Google Play and approximately 34,900 is the App Store

Capterra and GetApp

Unjudged. No listing was located under this product name on the date read, so these surfaces are named as unjudged rather than scored as zero

-

-

Gartner peer review platform

Unjudged. No product profile was located on the date read, so this surface is named as unjudged rather than scored as zero

-

-

Reddit, r/ManusOfficial

Not a rating and not scored. The community exists and its threads are visible through search indexing

-


Four honest notes on that table. First, the two figures that carry the most weight for a paid software decision are the two lowest, and they are the two with the smallest and the most complaint-shaped samples: G2 at 2.7 from 7 reviews, and Trustpilot at 1.2 across a page dominated by billing and refund complaints. Second, the mobile app store figures and the web review figures do not measure the same thing, because an app store rating is collected at the moment a person is using the product and a review-site rating is collected after something has gone wrong. Reporting both without that caveat would be misleading in one direction or the other. Third, G2 and Trustpilot both decline a plain automated read, and the figures above come from search-indexed retrieval of their public pages plus two independent aggregators that quote them with dates; no number in this table comes from a page this review could read directly, and each is attributed to the surface that carries it. Fourth, one aggregator publishes a large three-source rating that is arithmetically dominated by Google Play, which is why the components are shown separately rather than as a single headline number.


What Users Praise


The praise is consistent and it clusters in one place: research-shaped and multi-item work where the user knows what they wanted. Independent 2026 reviews describe Wide Research as the product's strongest differentiator and as a genuinely differentiated capability rather than a marketing layer. Reviewers with defined tasks report satisfaction, and the specific things named are the depth of unattended work, the breadth of artefact types from a single brief, and the fact that the agent keeps working while the user does something else. One technical overview published on arXiv describes the product as an early example of an agent that moves from intention to executed outcome rather than from question to answer. None of that is surprising, and it matches this review's Time and Quantity sub-scores, which are the two high ones.


What Users Complain About


The complaints are equally consistent, and they cluster on cost and support rather than on capability.


  • Credit consumption is unpredictable and arrives after the fact. This is the dominant theme across every platform. Users report entire monthly allocations consumed by a single task, tasks that loop and keep consuming, and no way to know the cost in advance. One review thread title from the product's own community reads as a cost complaint, and the independent coverage of the credit model describes the same thing in the same terms. A 2026 comparison of five agent platforms attributes the product's low G2 aggregate to this, reporting that roughly 42 per cent of its reviews are one-star ratings driven mostly by unpredictable credit consumption.

  • Billing, refunds and cancellation. Trustpilot is dominated by these. The complaints describe charges after a trial, refunds disputed, tickets closed without resolution, and support that feels automated. This review does not adjudicate those disputes, and it records them because a reader deciding whether to put a card on file should know what the review page looks like.

  • Reliability on long and complex tasks. Users report that long tasks need supervision, that quality falls off as a task extends, and that the product is weaker than expected on anything needing multi-session context. Independent reviewers summarise the split as satisfaction among research and automation users and frustration among those expecting cheap, predictable, polished output.

  • Output polish on builds. Reviewers consistently describe the application and website builds as fast prototypes rather than production-ready work, and the design control as below that of a dedicated tool.


Sentiment Summary


Overall sentiment: Sharply split, and split along a predictable line. The product is well regarded by users whose tasks are research-shaped, multi-item and self-verifying, and poorly regarded by users whose tasks are transactional, cost-sensitive or need polished output, and by users who have had a billing dispute. The app store aggregates are high, the professional review platforms are low, and both readings are real.


Key themes:


  • Cost opacity is the single most consistent complaint across platforms and it appears in every independent analysis. It is the same issue this review's Time sub-score reasoning names.

  • Reliability on long tasks is a repeated complaint, and no independent measurement of it exists in either direction. This is the absence that caps the Quality sub-score.

  • The billing and support complaints are a customer-service finding rather than a capability finding, and they belong in a procurement decision for an institution, because a disputed invoice with no resolution path is an operational risk rather than a quality one.

  • Praise concentrates where the product is differentiated and complaints concentrate where it is not, which is the same shape this review's scoring produces.


U365 Editorial Note


The crowd and the framework agree more closely here than on most tools in this series, and the point of agreement is the cost model.


Where they converge: the framework's Time sub-score is held below 8 because the spend is metered and unstated in advance, and the crowd's dominant complaint is the same fact experienced as an invoice. Both readings say the same thing, that the price of a run is disclosed after the fact. The framework's Skill sub-score and the Skill Illusion rating exist because the output arrives finished for a user who may not be able to judge it, and the crowd's frustration with polished-looking output that does not hold up is the same phenomenon seen from the other side. The framework's Quality reasoning names the absence of independent measurement; the crowd supplies the closest available substitute, which is a large volume of user reports about long-task reliability, and those reports are complaints rather than measurements.


Where the crowd is more useful than the framework can be: the volume of billing and refund complaints is something a review can only report and cannot adjudicate, and it is genuinely relevant to a U365 reader deciding whether to put an institutional card on a shared account. The framework has no dimension for vendor support quality, and the crowd is the only source on it.


Where the framework is more useful than the crowd: no review platform treats the standing instruction as a governance question. None connects a Skill built from a completed interaction, or a project master instruction applied to every future task, to the fact that the user now holds a documented capability they did not author. Directory sentiment has nothing to say about that, and it is the reason this review rates Skill Illusion High and takes Skill down to 3 while the same product sits at 4.6 across more than four hundred thousand app store ratings. That divergence is the case the framework exists to cover, and it is the most useful thing in this section for a U365 reader.


Manus: the crowd's platform ratings against the CI-First sub-scores, and the divergence this review records in Section 9

The one place the two genuinely disagree is the headline number. A reader who looks only at the app store aggregate sees a product people like. A reader who looks at Trustpilot sees one people are angry at. Both are measuring something real and neither is measuring CI-First benefit, which is why this review's 5.3 is not a reconciliation of the two but a separate judgement built from the framework.




Back to the TOC

Comparison and Alternatives


Alternative

"Choose the alternative if..."

"Choose Manus if..."

ChatGPT Agent and equivalents from the frontier model vendors (https://chatgpt.com/)

You already pay for a subscription to one of the frontier vendors and your tasks are mostly single-artefact and interactive. Agent mode inside a chat subscription has no separate credit meter, the cost is a known monthly figure, and the conversational loop is stronger for work you want to think through with a model rather than hand off

You want the task to leave the conversation: unattended, on a schedule, up to twenty at once, with a file system, an installable toolchain and a browser. The two are complements rather than substitutes, and the honest split is one-off thinking in the chat tool and list-shaped work on the agent platform

You want comparable one-prompt, walk-away autonomy at a published price you can plan around. Independent 2026 comparisons name Genspark as the closest all-purpose substitute, and it publishes its tiers in a form a reader can read rather than rendering them in a dynamic element

You need breadth of finished artefact types rather than one class of output, or the documented many-item mode. The reason to choose Manus is architectural: one dedicated agent and one fresh context per item is a different design from a single agent working a list

Cost predictability and source-cited output matter more to you than breadth of artefact. Its documents, slides, sheets and webpages agents share a research engine, it publishes a free tier and a per-month price, and independent coverage reports it winning on cost against this product

You want the widest artefact range from one subscription, or the scheduled and concurrent execution limits this product states. The honest caution for both is the same: independent coverage of the adjacent tool reports its Trustpilot rating at the same low level, with the same complaint pattern about trials, cancellations and credits, so the substitution is not a clean escape from cost disputes

Your need is recurring administrative work rather than research or building. Lindy is built around scheduled and triggered workflow execution, and independent 2026 comparisons report the strongest review track record among the alternatives

Your work is research, analysis, prototyping or content production at volume rather than a repeating administrative process. A workflow automation tool and an autonomous agent platform solve different problems, and a reader should be clear about which one they have

OpenManus, and other open agent frameworks (https://github.com/FoundationAgents/OpenManus)

You are a developer with cost sensitivity, a private environment requirement, and the appetite to maintain the agent yourself. You pay for model tokens rather than credits, and no task data leaves your infrastructure

You want the agent to be someone else's maintained product: the sandbox, the browser, the skills library, the connectors, the scheduled execution and the API, rather than a stack of components you assemble and keep running. The trade in the other direction is that a hosted agent means a hosted data relationship, which Section 7c sets out

A research service or a specialist freelancer

The output has to be right rather than drafted, the deadline is fixed, and a person is accountable for the answer. Any professional service you would otherwise commission is the comparison that matters most against this product

You want the first pass, you can judge it, and the cost of a human doing the assembly is the thing you are trying to remove. The two are not mutually exclusive: the most common serious pattern is using the agent to prepare the material and a specialist to make the judgement


Where Manus is clearly better. Two things. First, the many-item architecture and the artefact range taken together: no other product in this field documents a per-item agent with its own context at a tested scale of 250 items, and the same subscription covers a report, a spreadsheet, a deployed site, an application and a batch of images. Second, the candour of the vendor's own documentation on how the product works, its cost drivers, and the conditions under which its main mode does not fit.


Where Manus is clearly worse. Cost predictability, and the gap is wide. There is no pre-run estimate and no spend cap, the vendor's pricing page renders its figures in a form a reader cannot confirm, the two review platforms that matter for a paid decision rate the product poorly and cite cost as the reason, and its terms place the entire output-review obligation on you while warranting nothing about accuracy. Against an open framework you give up cost control and your own data custody. Against a chat subscription you give up a known monthly figure. Against Genspark and Skywork you give up a readable price. Against a professional service you give up accountability for the answer. Those are real trade-offs for a real capability, and a reader who needs predictability more than output volume should choose differently.




Back to the TOC

Verdict and Next Steps


Who should adopt it: A person or team with a recurring, list-shaped or assembly-heavy workload whose subject matter they can judge themselves, at a standard where they will read the artefact before it is used. The strongest fit is research, competitive intelligence, recurring reporting and first-pass building where you already know what a good output looks like because you have produced one by hand before.


When: The evaluation can start today on the free tier, and the honest first step costs nothing: run one list task of ten to twenty items with a named output format and read what comes back. Adopt, meaning subscribe, only after you have measured two things on your own work. The first is the credits your typical task consumes, because that is the number your subscription decision actually needs and only you can collect it. The second is whether you can tell a good output from a bad one on a task you have not done by hand before, because a person who cannot has bought an artefact-generator rather than a capability.


For what: The primary task is many-item research and first-pass production where the deliverable is a file. The secondary task is a recurring report or analysis you already know how to write. The task to avoid is anything you cannot evaluate, anything whose cost must be predictable, and anything you would send to a human without reading.


What to put in place before you subscribe: a written list of the task classes you will never hand over, a habit of reading the replay for any consequential run, an explicit rule that you write every project master instruction yourself, and a rule that no Skill enters your library without you reading it first.


UP-Context prompt pack:


UP-Context prompt pack. Three reusable prompts written in the U365 prompting method, each in the UP-Context order with a Profile line that assigns the CI-First Profile before the task and an output format that states what you do with the result. Copy them into Manus with your own context. The first is the boundary prompt to run once before you delegate anything real, and the third is the one that keeps the library honest over time.


Prompt pack 1: The boundary, stated before the first task.


Context: I am evaluating this platform for [the kind of work]. The tasks I have in mind are [name two]. The output has to fit [describe where it goes and who reads it]. I can personally judge [say what you can judge] and I cannot judge [say what you cannot].
Profile: Act as a Co-Worker and Assistant (level 2). I own what counts as done. You execute.
Task: Before we start, tell me back the boundary for this account: which of the tasks I described you are well suited to, which you are not, and what you would need from me that I have not given.
Constraints: Do not begin work beyond the boundary. Do not fill a gap in my brief with an assumption you do not state. If a requirement is ambiguous, name the ambiguity rather than choosing a reading silently.
Output format: Three short lists: suited, not suited, and what you need from me.

Prompt pack 2: The many-item research table, with the honesty constraint that matters most.


Context: I need a comparison of [n] items in the attached list, for [audience and decision]. The fields are [name four to six columns precisely]. I already know [state what you know] and do not need it re-established.
Profile: Act as an Analyst and Tester (level 4). You gather and flag. I interpret and decide.
Task: Research each item and produce one row each in a single table with the named columns. Cite a source for every cell and state the date you read it. Where you cannot establish a cell from a source, write "not established" rather than inferring it.
Constraints: Do not fill a cell by analogy. Do not merge two items into one row. Do not present a source you did not open. Prefer the item's own primary material over an aggregator. Flag anything you inferred rather than read.
Output format: The table, with a source column and a date column. Then a list of the items where two or more fields could not be established, with the reason.

Prompt pack 3: The library and standing-instruction audit, which is the prompt this review recommends running on a schedule.


Context: This account holds [n] skills and [n] project workspaces with master instructions, created over [period]. I am auditing what my account now instructs, not asking for new capability.
Profile: Act as an Analyst and Tester (level 4). I am checking my own standing instructions, and I wrote some of them and you wrote others.
Task: List every skill in my library and every project master instruction, and for each one state what it tells you to do, whether it was written by me or by you, and whether it contains an instruction that would apply to work outside the task it came from.
Constraints: Do not edit, create, delete or improve anything in this pass. Do not summarise; list. Where you cannot determine who wrote an artefact, say so rather than guessing.
Output format: Two tables, one for skills and one for project instructions, with the name, the instruction, the author, and your answer on scope. Then one line naming the artefacts you could not assess.
Read both tables yourself. Delete or rewrite what you no longer agree with, and keep the master instruction of every active project in your own words.

Related U365 content:


  • URC's INSIDE Tools Review of Klarent (fore ai), for the agent-platform case where a finished artefact arrives from a system that cannot be asked whether it asserted the right thing, and for the contrast with a platform whose artefacts are test results rather than documents.

  • URC's INSIDE Tools Review of Rabbit OS3, for the other Humics-Risky badge in this series and for the case where a vendor's permission model rather than its contract was the operative fact.

  • URC's INSIDE Tools Review of MiMo-V2.6-Flash, for the clearest precedent on scoring Quality down for the absence of independent measurement, and for the supplier-conduct section this review adapts to contract terms.

  • URC's INSIDE Tools Review of Jev AI (TypeSafe AI), for the decision-model case where the deliverable is a judgement rather than an artefact.

  • Browse the published U365 Tools Reviews index at https://www.university-365.com/tools


Integration surfaces, U.Copilot and SL-OS


Mastering this product means deciding where the agent's outputs and the agent's standing instructions live, and two U365 surfaces take that decision.


U.Copilot is the front door to the U365 tool library at https://www.university-365.com/ucopilot. Use it to design the workflow before you delegate anything, because the specification work is where this class of platform succeeds or fails and a prompt precise enough for U.Copilot to structure is usually precise enough for the agent to execute. Ask it to produce your task-class boundary list, to write the output format your organisation actually uses, and to name the verification step for each task class. Ask it to record which tasks you will never hand over.


SL-OS is where the results belong. Put the artefacts in the same places your other work lives rather than only inside the tool's workspace, and put the decisions about them in your LIPS Digital Second Brain under the right category, because the run record is the vendor's and the decision record is yours. The scheduled-task surface fits a review cadence rather than a daily routine, and the ULM dimensions it touches are Career and Quality of Life: the case for the tool is time returned from assembly, and the case against it is the same time spent judging output instead. Character is touched in one specific way, in refusing to ship an artefact you cannot defend.




Back to the TOC

Status and Last Tested


Re-check: trigger-based, maximum six months. The triggers are listed under Status and Re-check above, and the first three, a pre-run credit estimate or hard budget stop, an independent measurement of task reliability, and any change to the memory and skills write path, would each change a sub-score rather than only the wording.


What Active means here. Active means current and recommended for the work this review describes: research-shaped deliverables, many-item processing, and first-pass builds that a person reads before anything depends on them. It does not mean the product is verified. No independent measurement of task reliability exists, the two review platforms that carry the most weight for a paid software decision rate it poorly on cost, and no surface states what a task will cost before it runs. A reader who needs a fixed cost or a measured guarantee should treat those as unavailable and plan around them.


Version reviewed: Manus, as documented at manus.im in September 2026




Back to the TOC

Tool to Skill to Credential


No credential chain is asserted for this tool, and this section records why rather than filling the gap with a plausible programme name. U365 does not assess a competency in agent supervision and procedure review today; it would sit in UIT if a programme were built for it. No credential chain attaches to the act of delegating a task to an agent at any level: the delegation act is not an assessable artefact, and chains attach only to instruction specification, verification, standing procedure governance and the costed decision.


Tool skill

U365 competency

Credential

Institute

Writing an instruction precise enough that an agent can execute it: context, profile, task, constraints and output format, with the scale stated

Specification of delegated work. This is the competency the tool changes, and it transfers to every agent platform

Business Analysis Professional (60 days, diploma), verified published 2026-09-25. Its published programme includes Business Analysis Foundations, Agile Requirements, Business Process Modeling and Business Benefits Realization

UIT (Technology, AI, Data Science), against the applied AI and agent-architecture material

Reading a finished artefact carefully enough to know whether it is right, and saying what the agent should have done differently

Evaluation of a deliverable against a standard the Fellow sets, which is the verification habit the framework depends on

Business Analysis Professional (60 days, diploma), verified published 2026-09-25, through its Business Benefits Realization and Business Analysis Foundations content

UIB (Business Management, Entrepreneurship)

Reading a project's standing instruction and a saved Skill as the operating procedure they are, and deciding whether to keep them

Governance of agent-authored procedure, which no U365 institute currently assesses

Confirm with academic team. No published U365 programme assesses the supervision of an autonomous agent or the governance of agent authored procedure. Recorded as a curriculum gap. the competency would sit in UIT if a programme were built for it; the nearest existing certificate territory assesses building an automation rather than commissioning one.

UIT (Technology, AI, Data Science), where the competency would sit if a programme were built for it

Costing a run before commissioning it, including what an unattended task consumes and what it costs to redo

Technology investment appraisal, applied to a metered service whose price is disclosed after the fact

Financial Analysis Specialist (30 days, diploma), verified published 2026-09-25. Its published programme includes Corporate Financial Statement, Financial Modeling, Forecasting Financial Statements and Data and Economic Modeling

UIB (Business Management, Entrepreneurship)


Access level, stated plainly. University 365 has three academic access levels: DISCOVERY, INSIDER and SUPERHUMAN. Specialized diplomas and certificates carry Basic, Foundation and Expert levels: DISCOVERY Fellows can enrol in Basic level programmes only, INSIDER Fellows in Basic and Foundation programmes, and SUPERHUMAN Fellows in all of them. University degree programmes carry a single Expert level and are open to SUPERHUMAN Fellows only. No degree chain is asserted against any row in the table above, and no claim is made that completing any programme listed there awards credit toward a degree or toward a named micro-credential. The catalogue does not expose credit transfer or a per programme access level, and this review asserts neither.


A curriculum gap, recorded rather than filled. U365 publishes no credential that assesses the review of an agent's own procedural memory, which is the durable artefact this product creates and the mechanism behind the High Skill Illusion rating. The honest entry is that none exists. Two candidates are consistent with this review's own recommendation that the judgement stays outside the tool: agent supervision and procedure review, and metered automation economics. Neither is a current programme and neither is presented as one. Two gaps are recorded rather than one: agent and procedure supervision, which belongs to UIT, and authorship integrity for agent generated communication, which belongs to UIC.


No credential chain is mapped for the tool's own output. The subjects of the artefacts an agent produces are not taught by producing them, which is why the Skill sub-score is 3 and why this section asserts no chain. A Fellow who needs a subject-matter competency should take the programme that teaches the subject.




Back to the TOC

Migration Path


Not applicable. Manus is Active and recommended for the work this review describes. No Migration Path section is required and none is included, and no migration plan has been built. A reader who is currently using the product and wants to reduce dependence on it has a different exercise available, which is the library and standing-instruction audit in the prompt pack above: it is the step that keeps the account yours rather than the vendor's.




Back to the TOC

U365's Recommendations to Learn More


Official learning resources



Video tutorials and channels


The three walkthroughs below were each resolved against the YouTube oEmbed endpoint before publication. They are third-party explanations of the product rather than vendor material, and no finding in this review rests on them. For a full run rather than an edited excerpt, the replay links the vendor publishes for its Wide Research examples are more instructive than any edited video, and the documentation index is the fastest route to those.





Written tutorials and deep-dive articles



Community and social



One honest note: the community discussion of this product is large and it is mostly cost-shaped. There is no independent practitioner community measuring its reliability, and the absence of that measurement is a finding of this review rather than a gap in the search.


Resources on X


Dedicated X channels:


The Manus product account on X, the handle the vendor's own site links to


That is the only dedicated channel for this product confirmed for the purposes of this review, and the review deliberately lists one rather than filling the section with handles it has not checked. For a reader following the category rather than one vendor, the accounts that consistently carry independent agent evaluation and the model-supply layer are https://x.com/ArtificialAnlys for independent model measurement, and the frontier vendors at https://x.com/OpenAI, https://x.com/AnthropicAI and https://x.com/GoogleDeepMind for the components this product routes across. On this product in particular, the account to watch first is the vendor's own, because a pre-run cost estimate, a spend cap or a published rate card would be announced there before it reached a documentation page.




Back to the TOC

Glossary


CI-First Benefit Score


The average of four dimensions, each scored 0 to 10: Time, Quantity, Quality, and Knowledge and Skill. It answers whether using the tool makes Co-Intelligence more profitable than Human Intelligence alone. Bands: 0 to 2.0 CI-First Negative, 2.1 to 4.0 CI-First Neutral, 4.1 to 6.0 CI-First Positive, 6.1 to 8.0 CI-First Strong, 8.1 to 10.0 CI-First Transformative. The score accounts for the overhead of prompting, supervising and verifying, not just the benefit the tool produces. Manus scores 5.3.


CI-First Profile


The role the AI plays in your working relationship. (level 1) Co-Creator and Thought Partner, (level 2) Co-Worker and Assistant, (level 3) Coach and Tutor, (level 4) Analyst and Tester, (level 5) Challenger and Devil's Advocate. Lower level numbers indicate higher AI autonomy. Assigning a profile before giving the AI a task is a core CI-First discipline. Manus is primarily a Co-Worker and Assistant (level 2), with Analyst and Tester (level 4) as a substantial secondary and Co-Creator and Thought Partner (level 1) narrowly.


Humics Protection Badge


A rating of whether a tool protects, leaves neutral, or erodes three human capabilities: Creativity, Critical Thinking, and Social Authenticity. Each is scored +1, 0, or -1, and the sum gives the badge. +2 to +3 is Humics-Friendly, -1 to +1 is Humics-Neutral, -2 to -3 is Humics-Risky. It measures whether the tool strengthens the human or contributes to AI Obesity. Manus is Humics-Risky at -2 / +3: Creativity neutral, Critical Thinking eroded, Social Authenticity eroded.


AI Imposture Risk


The likelihood that a tool traps you in one of three illusions. The Time Illusion is the appearance of saving time when net time is lost. The Quantity Illusion is high volume that looks good but does not survive inspection. The Skill Illusion is the appearance of competence in you while the underlying skill is absent or eroding. Each trap is rated Low, Medium, or High with cited evidence, and the overall level is Low when all three are Low, High when two or more are High. Manus is Medium overall, with Skill Illusion High, which is the framework's clause 5.2.3-a floor in operation.


Collaboration Mode


How you and the AI divide the work. In Centaur mode there is a clear division of labour: you handle the judgement, the strategy and the final decision, and the AI handles the heavy drafting and processing. In Cyborg mode the two of you iterate rapidly with no clear boundary, which requires a stopping criterion you apply yourself and which is available only when the Imposture Risk is Low. Manus is Centaur, and Cyborg is not available on it.


Standing instruction


Any durable artefact that tells the agent what to do in future runs rather than in the current one. On Manus there are two: a project's master instruction, which applies to every task created inside that project, and a Skill in your library, which the agent loads on demand. Both outlive the task that produced them, which is why they carry the Skill Illusion risk described above.


User Sentiment


The aggregated public opinion from review platforms, community forums, and app stores. It is reported separately from the CI-First score because crowd sentiment can contradict a rigorous evaluation. Where the two agree, the finding is stronger. Where they diverge, the divergence is worth explaining. For Manus the crowd is sharply split: high app store aggregates, low professional review platform scores, and a consistent complaint about cost opacity that this review's Time reasoning independently reaches.


Review Status


Review Status records the current standing of the tool at the time of the last test. Active: the tool is current and recommended. Active (updated): recently re-checked and the content was refreshed. Changed: a re-check trigger fired and an update is pending, so read the review with that in mind. Risky: the tool has significant unresolved issues, or it has been clearly surpassed by newer alternatives. Use it with caution and read the Limits section. Stale: this review has not been re-checked in over 6 months, so treat details such as pricing and features as unverified. Retired: the tool still works but is no longer recommended. Deprecated: the tool has been shut down or fundamentally changed. Retired and Deprecated posts include a Migration Path section. Manus is Active, with the conditions stated at the status badge.




Back to the TOC

Sources


Vendor primary sources


  • Butterfly Effect Pte. Ltd., Manus welcome documentation, for the product's own definition of itself as an autonomous general agent, the sandbox description, and the claim of production-ready results: https://manus.im/docs/introduction/welcome

  • Butterfly Effect Pte. Ltd., plans and pricing documentation, for the credit-based model, the statement that consumption depends on task complexity, the three plan descriptions, credit refills and add-ons, and the statement that add-on credits never expire: https://manus.im/docs/introduction/plans

  • Butterfly Effect Pte. Ltd., pricing page, for the three credit allowances at 4,000, 8,000 and 40,000 credits per month, the 300 daily refresh credits, the annual saving of 17 per cent, the 20 concurrent and 20 scheduled tasks, and the Wide Research scaling statements: https://manus.im/pricing

  • Butterfly Effect Pte. Ltd., what are credits help page, for the three consumption drivers, the statement that credits are only consumed during active task processing, the full-refund policy for tasks that fail for technical reasons, the three worked usage examples with their durations and credit figures, and the credit expiry and consumption-order rules: https://manus.im/help/credits

  • Butterfly Effect Pte. Ltd., Team plan page, for per-seat billing, the shared credit pool, the 17 per cent annual saving, pooled-credit administration by owners and super admins, and the upgrade path from an individual plan: https://manus.im/team

  • Butterfly Effect Pte. Ltd., Manus Skills documentation, for the file and folder based skill model, the four routes into the library including building a skill from a successful interaction, slash-command activation, progressive disclosure, and the warning that community skills can contain code and shell commands: https://manus.im/docs/features/skills

  • Butterfly Effect Pte. Ltd., Projects documentation, for the master instruction and knowledge base model, availability on all tiers, private-by-default behaviour, the non-retroactive configuration rule, and the team visibility rules: https://manus.im/docs/features/projects

  • Butterfly Effect Pte. Ltd., Wide Research documentation, for the per-item agent architecture, the context window problem and the fabrication threshold at 8 to 10 items, the result synthesis description, the tested scale of 250 items, the comparison table against chatbots, and the stated conditions under which the mode does not fit: https://manus.im/docs/features/wide-research

  • Butterfly Effect Pte. Ltd., documentation index, for the product's surface map used to structure Sections 1 and 4: https://manus.im/docs/llms.txt

  • Butterfly Effect Pte. Ltd., Manus API documentation, for programmatic access to the agent: https://manus.im/docs/integrations/manus-api

  • Butterfly Effect Pte. Ltd., Terms of Use, last updated 28 November 2025, for the contracting entity Butterfly Effect Pte. Ltd., the artificial intelligence disclaimer and user responsibilities at section 1, the scope of the licensed rights over Your Content and the statement that the Company does not own Input or Output, the retention of Usage Data and of the skills, expertise and methods used to provide the service, the Master Services Agreement hierarchy at section 2.5 and 2.6, and the watermark statement: https://manus.im/terms

  • Butterfly Effect Pte. Ltd., Privacy Policy, last updated 21 August 2026, for the data processor and controller roles, the information collected including inputs, prompts, derived data and Space end-user data, the automatic collection section covering device data, sandbox environment data including files, shell commands, generated code and execution logs, and browser operator data including page content and actions performed on the user's behalf, the purpose of creating aggregated de-identified data and sharing it with third parties, the research and development lawful basis, the international transfer and European users notice, the Singapore incorporation and data protection officer statement, and the team privacy controls: https://manus.im/privacy

  • Butterfly Effect Pte. Ltd., Master Services Agreement page, which governs the Team plan and the API. The page returned a title and no readable body on the date read, and nothing in this review rests on its content: https://manus.im/policies/services-agreement

  • Butterfly Effect Pte. Ltd., independent operations notice, for the vendor's own account of resuming independent operation, the data backup and restoration guidance, and the statement about a temporary interruption to access for some users: https://manus.im/blog/manus-resumes-independent-operations


Independent sources


  • Minjie Shen, Yanshu Li, Lulu Chen and Qikai Yang, "From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent", arXiv 2505.02024, an independent academic overview, for the architecture description, the application survey and the authors' statement that benchmark results are reported rather than measured by them: https://arxiv.org/abs/2505.02024

  • Codebasechat, 2026 review, for the platform-by-platform rating table with sample sizes, including the G2 figure of 2.7 from 7 reviews, the Trustpilot figure of 1.2 from approximately 160, the Product Hunt figure of 4.4 from 9, and the finding that the product is not a coding assistant substitute because it holds no multi-repository memory across sessions: https://codebasechat.com/review/manus-review

  • Delta4, 2026 review, for the two-task hands-on methodology, the credit-consumption findings, and the comparison to adjacent tools including the report that Skywork carries the same complaint pattern on Trustpilot: https://delta4.io/blog/manus-ai-review

  • PlugThis, 2026 review for creators, for the account of credit consumption in the hundreds of credits per task and the absence of a spend cap, and the description of the browser tooling as designed for personal use: https://plugthis.ai/blog/is-manus-worth-it-in-2026-an-honest-review-for-creators

  • Superapp, 2026 review, for the split-sentiment table by dimension, including the assessment that autonomous research and repeatable automation are strong while cost predictability is poor and long-task reliability is inconsistent: https://www.superappp.com/blog/manus-ai-review-2026-pricing-reviews-is-it-worth-it

  • Lindy, 2026 review, written by a competing vendor and read as such, for the description of the multi-agent architecture and the assessment of Wide Research as a genuine technical differentiation rather than a marketing layer: https://www.lindy.ai/blog/manus-ai-review

  • a 2026 five platform comparison by an industry publication, for the aggregate statement that the low G2 rating is driven mostly by unpredictable credit consumption and for the comparison against Genspark, Taskade, Skywork, Lindy and chat agent mode: https://aistartupinsights.com/compare/manus-alternatives-5-real-ai-agents-compared-for-2026

  • SearchTools.ai, aggregator entry, for the three-source rating reading of 4.60 across approximately 443,400 reviews with the Google Play and App Store components shown separately: https://searchtools.ai/t/manus

  • AgentsIndex, product entry, for the dated quotation of the G2 and Trustpilot figures and for the review themes on interface ease, pricing and credit restrictions: https://agentsindex.ai/manus-ai

  • TechReviewer, 2026 review, for the review-corpus analysis based on 15 reviews across two platforms between March 2025 and June 2026, including the finding that approximately a third of the corpus predates the reviewed period: https://techreviewer.co/products/manus

  • Google Play, Manus app listing, for the app store rating and count as reported by the store page: https://play.google.com/store/apps/details?id=tech.butterfly.app

  • People of Internet, research write-up of the regulatory decision, for the chronology of the acquisition and its unwind, the grounds stated in the decision, and the data isolation and deletion requirement. Read as published reporting, not as a verified fact of this review: https://peopleofinternet.com/articles/beijing-s-forced-unwind-of-the-meta-manus-deal-extends-inves.html


Readability of the sources above, stated rather than hidden. Every link in this review was checked on 2026-09-25, and the product's X account was confirmed by reading the profile rather than by assuming the handle. G2, Trustpilot, Capterra, GetApp, Product Hunt and Reddit do not serve their content to a plain automated read, so the figures attributed to them in Section 9 come from search-indexed retrieval and from independent aggregators that quote them with dates, and each figure is attributed to the surface that carries it rather than to this review. No rating or review count in this review comes from a page this review could not read directly. The vendor's pricing page renders its currency figures in a dynamic element, which is why the currency figures in this review are attributed to two independent reviews and the credit allowances are attributed to the vendor's own page, and why the review does not state a currency price it could not confirm. The Master Services Agreement page returned a title and no readable body, and nothing here rests on it. The three video walkthroughs in the recommendations section were each resolved against the YouTube oEmbed endpoint before publication.


Community and community-reported evidence



Internal sources


  • CI-First Evaluation Framework v1.2, published at https://www.university-365.com/tools, the scoring rubrics in Section 3, the Humics rating and clause 4.2-a in Section 4, the imposture risk assessment and clause 5.2.3-a in Section 5, the profiles in Section 6, the collaboration modes and clause 7.5 in Section 7, and the scoring procedure and principles in Section 9.

  • INSIDE Tools Post Template, including the Agent Platform variant, which applies here: https://www.university-365.com/tools

  • Published INSIDE Tools Reviews used as internal comparisons, listed at https://www.university-365.com/tools: Klarent (6.0, the most recent agent-platform application and the source of this review's Section 7c contract-terms method), Rabbit OS3 (4.5, Humics-Risky, the other application of clause 4.2-a on a permission surface), MiMo-V2.6-Flash (5.5, the supplier-conduct precedent and the clearest precedent for scoring Quality down for unverifiable measurement), MiMo-V2.6-Pro (5.8), Jev AI (5.5) and Hermes Agent (7.8, the highest-scoring agent platform in this series).




Back to the TOC

Faculty Note on Evidence Quality


Eight items did not survive checking, and the pattern across them is worth stating before the list. This is a vendor whose documentation is careful, specific and at times self-critical, and whose marketing surfaces do not agree with its own documentation or with each other.


First, the marketing claim and the terms are in direct conflict about what you receive. The pricing page and the Team page describe the product as business artificial intelligence that "works like your best employee", and the documentation states that it delivers "production-ready results without you managing every detail". The Terms of Use state the opposite in the operative clauses: that AI systems "are based on probabilistic models, which may result in misunderstandings or errors", that the Company "is not responsible for any misunderstandings or inaccuracies caused by AI", that you "are responsible for independently reviewing all Output", that you are "fully responsible for monitoring and approving the use of Output", and, in the disclaimer, that there is no warranty that the output "WILL BE ACCURATE OR RELIABLE". Both sets of words are the vendor's. A reader should treat the documentation and the terms as the description of the product and the marketing sentence as a hypothesis.


Second, the Wide Research comparison table sets the product against an unnamed weaker competitor rather than against a measured baseline. The table's left column is "AI Chatbot" and its right column is this product, with the chatbot characterised as degrading beyond 8 to 10 items and this product as offering uniform quality at any scale. The chatbot column is presented as a general fact with no citation, the product column is presented as a capability with no measurement, and the section header states that "No other AI tool can handle this scale", which is a comparative claim against the entire field with no comparison performed. The claim that item 250 receives the same depth as item 1 is a design property, which is defensible, and it is presented in the same register as the performance claims, which are not measured.


Third, the scale claim moves within the vendor's own page. The body of the Wide Research page states that the mode has been tested up to 250 items and that the limit is theoretically unlimited in practice depending on task complexity. The comparison table states that it scales to hundreds seamlessly. A reader planning a job needs the tested figure, not the theoretical one, and the two are presented as though they were the same claim.


Fourth, the entry price is not confirmable on any vendor surface. The pricing page's three plan cards render their currency figures in a dynamic element, so a reader can confirm the credit allowances and cannot confirm the price. Two independent 2026 reviews report $20 per month for the 4,000-credit card and they agree with each other, and the vendor publishes no readable figure to check them against. The 17 per cent annual saving is stated on the pricing page and on the Team page and cannot be checked against a readable monthly list price, which is the promotional-base problem the framework's Pitfall 20 describes: a discount measured against a base the reader cannot see is a marketing figure rather than a price. What a reader can verify is the credit allowance, and the credit allowance is the number a budget should be built on.


Fifth, the plan structure differs between the vendor's own two surfaces. The documentation describes three plans named Free, Pro and Team. The pricing page presents three cards with no names, labelled by usage intent, and does not present a Team card. The Team plan lives on its own page. A reader comparing plans across those two documents is looking at two different taxonomies of the same product.


Sixth, the Team plan figures in circulation are not vendor-confirmed. Independent coverage reports the Team plan at $39 per seat per month with a five-seat minimum, a shared pool of 19,500 credits and a $195 monthly total. The vendor's Team page states per-seat billing, a shared pool, add-on credits and a 17 per cent annual saving, and does not render a figure. The independent numbers are consistent with each other and they are not vendor-confirmed, and this review reports them as reported figures rather than as a rate card.


Seventh, the benchmark result that circulates most widely has no run report behind it. A figure of 86.5 per cent at level one of the GAIA benchmark is repeated across aggregator sites and appears in summaries of the product, and an independent academic overview states that state-of-the-art results were reported on GAIA. Neither carries a dated vendor run report with the configuration, the evaluation setup, the number of attempts per question or the comparison set, and the aggregator sites carrying the figure are promotional rather than evaluative. This review therefore does not use the figure, and it scores Quality down for the absence of a published measurement rather than for any observed failure. That is the reason Quality is 5 and not higher, and the band at that sub-score is Moderate, which should not be read as a rejection of the product.


Eighth, the corporate transition is described at two very different strengths. The vendor's own notice describes resuming independent operations, thanks users for their patience, and says that for some users the transition "required backing up and restoring data, as well as navigating a temporary interruption to access". Independent press reporting describes a completed acquisition ordered unwound by a national regulator under a foreign-investment security review, with data isolation and deletion required and a reported deal value above two billion dollars. The vendor's account is accurate as far as it goes and it is not a contradiction in the way the first item is, so this is recorded as a difference in emphasis rather than as a finding. It matters to a university reader because it bears on where a data-processing relationship sits and how stable it is, and because a reader building a research corpus on the platform should know the history before they build one.


What the vendor got right, stated with the same emphasis. The Wide Research page documents the context-degradation problem honestly, including the fabrication threshold and the per-run quality curve, and it states the four conditions under which the mode is not the right tool. The credits page gives three worked examples with durations and credit figures rather than a vague assurance about fairness, and states the refund position for technical failures. The Skills page warns that community skills can contain code and shell commands and offers to audit them. The browser extension documentation states exactly what the agent can read and do, including that it reaches content from premium and authenticated services, and states the revocation path. The Projects page states that configuration changes do not apply retroactively. The terms do not claim ownership of your input or output and say so plainly. And the terms state the limitations of the technology, including that outputs may contain errors and that artificial intelligence lacks creative thinking and cannot express emotions as humans do, in language considerably more direct than the category norm. That is a vendor documenting its own product's hard edges, and it is the most useful material on the site for a reader deciding how to use it.


The pattern is consistent and it is the reason this review's score sits where it does. Where the vendor is describing a mechanism, a limit or a cost driver, it is specific and often self-critical, and that material carries the review. Where it is describing a result, the numbers are either unmeasured or unreadable, and that gap is what holds Quality at 5 and Skill at 3 rather than in the Strong band.


Review conducted by URC under the CI-First Evaluation Framework, version 1.2. Scoring date 2026-09-25. Tool version reviewed: Manus, as documented at manus.im in September 2026. Framework version applied: 1.2. Framework clauses checked: 5.2.3-a applies and Skill Illusion is High, on the mechanism of agent-authored skills built from completed interactions and project master instructions carried into every future task; 4.2-a applies on its first condition and Social Authenticity erodes, on the mechanism of an agent that acts and composes under the user's own authenticated identity in the user's own mail and logged-in sessions; 7.5 returns a null, because the several agents this product runs operate inside one execution under one orchestrator rather than in a shared channel with named participants. Collaboration Mode: Centaur, required, with Cyborg unavailable.


CI-First Evaluation Summary Card


Field

Value

Tool

Manus AI

Vendor

Butterfly Effect Pte. Ltd., Singapore

Category

Applied AI / Agent Platform

Version reviewed

Manus, as documented at manus.im in September 2026

Framework applied

CI-First Evaluation Framework v1.2

Time

7 / 10 (Strong)

Quantity

6 / 10 (Moderate)

Quality

5 / 10 (Moderate, capped by the absence of independent measurement)

Skill

3 / 10 (Marginal, and the Skill Illusion rating is High under clause 5.2.3-a)

CI-First Benefit Score

5.3 / 10

Band

CI-First Positive (4.1 to 6.0)

Humics

Creativity 0, Critical Thinking -1, Social Authenticity -1

Humics Badge

Humics-Risky (-2 / +3)

Imposture Risk

Time Medium, Quantity Medium, Skill High

Imposture Level

Medium overall

CI-First Profile

Primary Co-Worker and Assistant (level 2); secondary Analyst and Tester (level 4); narrowly Co-Creator and Thought Partner (level 1)

Collaboration Mode

Centaur, required. Cyborg not available

Clause 5.2.3-a

Applies. Skill Illusion High

Clause 4.2-a

Applies on its first condition. Social Authenticity erodes

Clause 7.5

Null. Several agents inside one execution, not a shared room

Status

Active

Last tested

2026-09-25

Re-check

Trigger-based, maximum 6 months

Independent measurement

None published in a methodologically documented form


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
Image by Erik  Lucatero

Become Superhuman

Master AI to stay irreplaceable in every field.

 

 

 

​

​

Apply for Admission Today.
Select Your Initial Access Level.


Become a DISCOVERY, INSIDER, or SUPERHUMAN Fellow.

Image by Milad Fakurian

Master Your Life with a Digital Second Brain

Turn overwhelm into clarity with LIPS + CARE
U365’s unique framework to organize your goals, projects, and knowledge into a superhuman system for success

bottom of page