Market map / Battle
Market map · scheduled battle

Browser agents and computer-use automation

September 13 at 20:42 · $0.091 total

Judged a tie
judged by grok-4.6 · passes split → scored as a tie
“Cleaner joints (labs vs browsers vs RPA vs infra vs verticals), fewer misplacements, sharper value-flow to foundation APIs, credible whitespace.”
claude-fable-5
$0.06321s1221 tokens

Market Map: Browser Agents & Computer-Use Automation

(1) Thesis

The market is restructuring from brittle, rules-based RPA sold to enterprises top-down into general-purpose vision/action agent models sold bottom-up as APIs and consumer products — collapsing the value of workflow-specific automation and shifting it to whoever controls the agent-capable foundation model and the browser/OS surface it acts on.

(2) Segments

A. Foundation model "computer use" APIs

  • Anthropic (Claude computer use), OpenAI (Operator / computer-using agent), Google DeepMind (Project Mariner)
  • Dynamics: Frontier labs commoditize the core capability (screen perception + action). Everyone downstream builds on or against these. Rapid capability gains, punishing for thin wrappers.

B. Agentic browsers & consumer agents

  • Perplexity (Comet), The Browser Company (Dia), Opera (Browser Operator), OpenAI (ChatGPT Agent as consumer surface)
  • Dynamics: Battle to own the end-user surface where agents run — distribution play against Chrome. High burn, uncertain monetization, but the winner gets default-agent status.

C. Enterprise agentic automation / RPA incumbents pivoting

  • UiPath, Automation Anywhere, Microsoft (Power Automate + Copilot Studio), Salesforce (Agentforce)
  • Dynamics: Incumbents retrofitting LLM agents onto governance, connectors, and existing contracts. Distribution and compliance moats vs. technical debt and cannibalization risk.

D. Developer infrastructure for browser agents

  • Browserbase, Browser Use, Anon (less sure on current traction), Anchor Browser (less sure), Steel.dev (less sure)
  • Dynamics: Picks-and-shovels: headless browser infra, session management, auth handling, anti-bot navigation. Fast-growing but exposed to labs bundling infra themselves.

E. Vertical agent applications

  • Harvey (legal), Sierra (customer service), 11x (less sure on health of business), Decagon (support)
  • Dynamics: Wrap agent capability in domain workflows, data, and outcome-based pricing. Defensibility comes from workflow depth and eval/guardrail investment, not the model.

(3) Value winner

Segment A (foundation labs) captures the most value: computer-use capability improves with general model scale, so labs get it "for free" while everyone else rents it; they can vertically integrate downward (Operator shows labs will own the consumer surface too) and their APIs tax every other segment. Second-most durable is Segment C — enterprise buyers pay for governance and integration, not raw capability — but incumbents may lose to Microsoft's bundling.

(4) White-space opportunities

  1. Agent identity, permissions & audit layer — a neutral "IAM for agents": scoped credentials, spend limits, action logs, and legal accountability when an agent transacts on a user's behalf. Fragments of this exist (Anon, Stripe's agent toolkit) but no one owns the trust/liability layer across sites.
  2. Agent-native website interfaces / access marketplace — infrastructure letting publishers and SaaS apps expose structured, monetized "agent lanes" (rate cards, verified access, structured actions) instead of the current adversarial bot-blocking arms race. Cloudflare is adjacent (pay-per-crawl) but no one covers transactional agent access end to end.
gpt-5.6-luna-pro
$0.00930s6287 tokens

1) Thesis

Browser agents are restructuring automation from deterministic API/workflow execution toward model-driven interaction with the existing graphical web, shifting value from task-specific bots toward reliable execution infrastructure, enterprise control planes, and proprietary workflow context.

2) Market map

Segment 1: Computer-use models and agent runtimes

Companies: OpenAI, Anthropic, Google, Browser Use, Adept

  • OpenAI — Operator and the Computer-Using Agent (CUA) model provide general-purpose visual browser/computer interaction.
  • Anthropic — Claude’s “computer use” capability enables screenshot-driven mouse and keyboard control.
  • Google — Project Mariner and related Gemini agent research target browser navigation and multi-step web tasks.
  • Browser Use — Open-source browser-agent framework that converts natural-language instructions into browser actions.
  • Adept — Pioneer in action-oriented foundation models for interacting with software; its technology and team were acquired in part by Amazon in 2024.

Dynamics: Model capability is improving rapidly, but raw benchmark performance is not enough; reliability, latency, permissioning, recovery from failure, and access to browser state are becoming the real differentiators. The model layer is likely to commoditize faster than the surrounding infrastructure and workflow data.

---

Segment 2: Browser execution and agent infrastructure

Companies: Browserbase, Steel, Browserless, Bright Data, Hyperbrowser [less certain on relative market traction]

  • Browserbase — Cloud browser infrastructure, session management, browser observability, and tooling for AI agents.
  • Steel — Open-source/cloud browser infrastructure focused on programmable browser sessions for agents and automation.
  • Browserless — Hosted headless-browser infrastructure, primarily for automation, scraping, and browser control.
  • Bright Data — Proxy, web-access, scraping, and browser-rendering infrastructure used in automated web interaction.
  • Hyperbrowser — Browser automation infrastructure positioned for AI-agent workloads. [Less certain: relatively early-stage compared with the companies above.]

Dynamics: This layer provides the scarce operational primitives—persistent sessions, CAPTCHA/proxy handling, isolation, authentication, recordings, replay, and scaling. It should benefit from every agent framework, although infrastructure may become partially bundled by hyperscalers and model providers.

---

Segment 3: Enterprise automation and orchestration platforms

Companies: UiPath, Automation Anywhere, Microsoft Power Automate, ServiceNow, Workato, Zapier

  • UiPath — RPA and enterprise automation platform expanding toward AI agents and unstructured computer interaction.
  • Automation Anywhere — Enterprise RPA platform adding generative AI and agentic automation capabilities.
  • Microsoft Power Automate — Large installed base across Microsoft 365, desktop automation, APIs, and emerging Copilot-driven workflows.
  • ServiceNow — Workflow platform using AI agents across IT, customer service, employee, and enterprise processes.
  • Workato — Integration and workflow automation platform increasingly incorporating AI agents and enterprise actions.
  • Zapier — Broad SMB/mid-market automation distribution, with AI agents layered onto its application-action ecosystem.

Dynamics: These firms own the enterprise distribution, permissions, connectors, audit trails, and workflow context that production deployments require. Their challenge is that traditional RPA is brittle, while fully autonomous browser agents remain difficult to govern; the likely product is a hybrid of APIs, deterministic workflows, and computer-use fallback.

---

Segment 4: End-user and vertical browser-agent applications

Companies: The Browser Company, Perplexity, MultiOn, Skyvern, Genspark [less certain on long-term traction]

  • The Browser Company — Developer of Dia, an AI-native browser positioned around browsing, research, and task completion.
  • Perplexity — Comet browser and agentic search/browsing initiatives aimed at making the browser an AI task environment. [Product availability and positioning are evolving.]
  • MultiOn — Consumer-oriented browser agents for shopping, booking, and other web tasks.
  • Skyvern — Visual/browser automation platform for workflows such as form completion, onboarding, and back-office operations.
  • Genspark — Agentic application platform that has offered browser/computer-use-style task execution. [Less certain: product scope and commercial durability are evolving.]

Dynamics: These products compete on task success and user experience rather than raw model intelligence. Consumer applications have large upside but difficult economics: browser sessions are expensive, websites change constantly, trust is low, and many high-value tasks require payments, identity, or liability.

3) Segment likely to capture the most value

Enterprise automation and orchestration platforms

This segment is best positioned to capture the largest durable share of value because it controls:

  1. Distribution — Existing relationships with CIOs, operations teams, and business-process owners.
  2. System access — Connectors, APIs, desktop environments, credentials, and permissions.
  3. Governance — Audit logs, approval flows, policy enforcement, security, and compliance.
  4. Workflow context — Knowledge of the business process, exceptions, service-level requirements, and downstream consequences.
  5. Hybrid execution — The ability to use APIs when available, deterministic automation when reliable, and browser/computer use only as a fallback.

The model and browser layers will likely be highly competitive and increasingly commoditized. A specialized agent application can win a use case, but the control plane that manages thousands of agents across an enterprise is more likely to become mission-critical and retain pricing power.

4) White-space opportunities

A. Transaction-safe “agent control plane” for consequential actions

A neutral platform that lets agents execute high-impact actions—payments, refunds, hiring decisions, account changes, regulatory submissions—with:

  • fine-grained permissions,
  • staged execution and human approval,
  • rollback or compensating transactions,
  • immutable evidence trails,
  • policy enforcement across browser and API actions,
  • insurance/risk scoring tied to task type.

Current platforms offer pieces of this, but there is not yet a broadly adopted, cross-agent standard for transaction integrity and liability when an agent acts through a graphical interface.

B. Continuous website/API adaptation layer

A shared system that monitors websites and enterprise applications for UI, authentication, and policy changes, then automatically:

  • detects broken agent workflows,
  • identifies the changed page element or business rule,
  • proposes and tests a repair,
  • validates the repair in a sandbox,
  • deploys versioned updates with rollback,
  • maintains a machine-readable “capability and reliability map” for each site.

Browser automation vendors provide selectors, recordings, and observability, but the market lacks a dominant self-healing, cross-agent maintenance layer that treats websites as continuously changing dependencies rather than static automation targets.