← Back to posts

Desktop Agents Are Not Bigger Chat Windows: What WorkBuddy and QoderWork Reveal

WorkBuddy and QoderWork are moving AI from answering questions to operating software and delivering files. The real contest is not feature count, but earning the right to take work off the user's hands.

Desktop agents became a recognizable product category in 2026. Tencent positions WorkBuddy as an all-scenario AI office workbench. Qoder launched QoderWork in February to extend its coding-agent capabilities into files, browsers, and everyday knowledge work. Both make the same promise: stop merely asking AI questions and hand it the job.

This is more than moving a chatbot into a desktop app. A chatbot's contract is to answer: the user asks, and the model returns language. A desktop agent's contract is to accept delegation: the user supplies an objective, context, and authority; the system plans and acts; a real file or software state is left for review. The second contract carries responsibility for process, result, and consequence.

A desktop agent is therefore a new transaction layer between human intent and software. People previously translated goals into menus, buttons, spreadsheets, and sequences across applications. The agent tries to absorb that translation and execution. It is not competing for screen space. It is competing to become the front door to work.

WorkBuddy's public product model resembles a configurable AI office. It separates Skills, Experts, and Expert Groups: a Skill is an ability, an Expert combines ability with experience, and an Expert Group adds roles and a collaboration process. Connectors, knowledge bases, automation, project workspaces, and cloud-hosted assistants turn the product into an organizational metaphor rather than a single assistant.

The official WorkBuddy interface showing tasks, experts, skills, connectors, knowledge bases, and automations
WorkBuddy puts experts, skills, connectors, knowledge, and automation inside one desktop workbench. Source: official WorkBuddy product page.

The important choice is not multi-agent architecture by itself. It is the intended user mindset. A user should not need to know how many models are running; they select a researcher, presentation specialist, or team and describe an outcome as if assigning work. WorkBuddy's public Ask, Plan, and Craft modes can also be read as an autonomy dial: return an answer, expose a plan, or produce the artifact.

QoderWork takes a different route. It anchors a task in a local folder workspace. New Task, Skills, schedules, messaging channels, and history live in the sidebar; conversation occupies the center; a Task Monitor appears during execution. The user sees a plan, tool calls, and file changes, not a cast of anthropomorphized digital colleagues.

This is a substantive distinction. WorkBuddy's product unit looks more like organizational capability; QoderWork's looks more like a task, workspace, and deliverable. One is naturally suited to capturing individual expertise as reusable experts and team processes. The other is naturally suited to moving continuously from research and organization to generation and revision inside a concrete project.

The official QoderWork interface showing a local folder workspace and file, content, and document tasks
QoderWork starts from a local folder and a concrete deliverable, not from an abstract agent architecture. Source: official QoderWork product page.

QoderWork's Design feature hints at the likely destination of the category: vertical workbenches growing above a general agent. Design does not merely return an image. It offers an infinite canvas, a structured plan, local editing, and handoff to Qoder IDE. The general agent interprets the assignment; a domain-native surface makes the artifact inspectable and editable.

On a checklist, the two products already converge: local files, browser automation, computer use, Skills, connectors, schedules, messaging, memory, and multiple models. As feature sets converge, however, experience becomes the harder differentiator. Users care less about whether an agent can click than whether it knows what to click, can recover from a mistake, and requires less supervision than doing the work manually.

The first challenge is task entry. A blank prompt is maximally flexible, but it transfers product-design work to the user. They must choose task size, attach context, restrict scope, and invent acceptance criteria. Experts, templates, folder workspaces, and vertical workbenches matter because they narrow the problem before the first prompt. Generality sells the demo; specificity produces the first success.

The second challenge is time. Chat products optimize for time to first token, while serious tasks may run for minutes or longer. Typing indicators are not enough. Users need a plan, current step, expected time and cost, the reason for a block, interrupt points, and recovery. The mature desktop-agent interface may look less like a conversation and more like an inbox, operations console, and review queue.

The third challenge is authority. QoderWork directly operates only in authorized directories and makes deletion recoverable. Its documentation also clarifies an important nuance: file operations happen locally, but text required by the model is sent to the selected LLM API provider. WorkBuddy offers different permission modes and asks for confirmation around higher-risk actions. Local-first is not a synonym for fully offline; data boundaries must be explained across files, apps, credentials, models, telemetry, and sync.

The fourth challenge is delivery. Users do not receive value because an agent prints “done.” They open the Word file, spreadsheet, deck, design, or reorganized directory and judge the numbers, format, sources, and editability. A better north-star metric than task starts or conversation turns is the amount of user supervision required per accepted deliverable.

The desktop agent delegation contract from brief, plan, and permission to action, artifact, and acceptance
A task becomes a real delegation only when goals, plans, permissions, execution, artifacts, and acceptance form one contract.

Together these surfaces form a delegation contract: is the objective clear, is the plan visible, is authority scoped, can execution be interrupted, is the artifact editable, and can the result be verified or rolled back? Conversation creates the contract. Retention depends on whether the product fulfills it.

This creates an autonomy paradox. A more capable agent does not automatically make a user feel lighter. If a ten-minute chore requires five minutes of briefing, three permission prompts, constant monitoring, and a final rewrite, manual work wins. Greater autonomy becomes a product advantage only when it reduces supervision while preserving control.

The user's mental model can progress through four levels. A question-answering tool tells me how. A copilot works with me. A contractor completes the job for my review. An operating system keeps the process running and calls me only for exceptions. WorkBuddy and QoderWork have entered the third level and are probing the fourth through schedules, messaging, and cloud execution. “Digital employee” remains a title that reliability must earn.

A reasonable counterargument is that desktop agents are simply easier RPA. That is half right. Both execute across applications, but RPA normally encodes a stable process in advance; an agent uses natural language to absorb variation and ambiguity. It lowers configuration cost while adding uncertainty. Reliable products will combine file interfaces, application APIs, connectors, and GUI operation. Computer use is more likely to become a compatibility layer for legacy software than the primary path.

Growth introduces a counterintuitive problem: when a product can do anything, users often do not know what to do first. “Complete any task from one prompt” may drive downloads but rarely creates a habit. The strongest wedges are frequent, disliked, relatively stable jobs with obvious deliverables—weekly reports, research collection, file organization, sales analysis, and decks with a standard format.

WorkBuddy's potential flywheel starts with Tencent's office ecosystem. Messaging, documents, meetings, email, knowledge, and identity reduce the cost of connecting work context. A successful assignment can become an Expert, Skill, or project process and then spread through a team workspace. Its larger opportunity is not merely personal productivity, but converting what one person knows into a capability an organization can repeatedly invoke.

Tencent's Q1 2026 results said the company believed WorkBuddy was already China's most widely used productivity-agent service. A June investor deck reported 80% next-day retention for active users and 88% for paid users across CodeBuddy and WorkBuddy combined. These are Tencent's claims, not independent market-share evidence, and the retention figures are not WorkBuddy-only. They do show that Tencent is using existing distribution, accounts, and subscription coordination aggressively.

QoderWork's growth path runs from Qoder's coding agent into adjacent work. Shared accounts and credits, local workspaces, Skills, connectors, and vertical workbenches such as Design can bring developers, designers, and knowledge workers into one execution surface. Public sources do not provide credible user or revenue figures for QoderWork. The interesting evidence is its capability-expansion path, not an invented growth race.

Both products will treat Skill marketplaces as leverage, but raw Skill count is likely to become a vanity metric. A valuable Skill is not a stored prompt. It is a verified work recipe with suitable inputs, tool calls, permission scope, output schema, failure handling, and acceptance checks. The flywheel begins when a successful task becomes a recipe, gets reused, and then spreads through a team.

A desktop agent growth flywheel from a useful first artifact to a saved skill, recurring delegation, and team distribution
Universal capability may earn a trial. A verified recurring job, reusable recipe, and team distribution earn retention.

This framing changes the metrics. Activation is the first useful artifact, not the first message. Retention is recurring delegation, not opening the client every day. Distribution is an artifact or Skill adopted by colleagues, not an invitation sent. Monetization should reflect the value of completed work, not merely token consumption.

Over the next 12 months, product shapes are likely to converge. Ask, Plan, Execute, and Review will become a common skeleton. Authorized workspaces, change previews, task history, recovery, schedules, and messaging entry points will become table stakes. Marketing will continue to sell universal agents and AI teams; recurring reports, research, organization, and presentation work will produce most retention.

Over 12 to 24 months, competition is likely to shift from capability demonstrations to reliable delivery. Skills will evolve from prompt bundles into executable recipes with versions, dependencies, permissions, and evaluation. Products will expose expected duration, cost, source coverage, and acceptance checks. Enterprise editions will center audit logs, credential custody, approval gates, data residency, and inherited identity. Models will remain important, but increasingly replaceable.

Beyond that, the desktop may be an execution location rather than the primary interface. A user initiates work from messaging, email, calendars, or a phone; a local client handles sensitive files and native applications; the cloud runs long jobs. QoderWork's current local scheduler cannot run when the computer is asleep or off. That limitation is evidence for an eventual hybrid of local execution and cloud continuity.

The likely destination is not an unbounded general agent with arbitrary control of an entire computer. Security, accountability, cost, and reliability favor constrained workspaces, explicit task contracts, and domain workbenches. Operating-system vendors and office platforms own distribution and permissions; independent products can still win through cross-ecosystem neutrality, deeper workflows, and a better trust experience.

Desktop-agent competition will not be settled by model count, the size of a fictional expert roster, or how humanly software moves a mouse. The winner must turn ambiguous language into a traceable plan, privileged execution into a recoverable process, and output into accepted work. Users are not merely handing over a task. They are handing over control, and the product has to earn it back with every delivery.

— End —