Browser MCP for AI Agents
Many web tasks can't be completed just by reading search results or downloading static HTML. An AI agent often needs to open a website, click a button, fill out a form, wait for JavaScript to execute, navigate across multiple pages, or securely maintain an authenticated session.
Browser MCP equips AI agents with fully featured browser automation tools directly through the Model Context Protocol.
By integrating the 2Captcha Browser MCP, an MCP-compatible agent can launch a managed browser to interact with dynamic elements, execute custom JavaScript, extract content, persist session cookies, and seamlessly handle heavily protected websites. Instead of forcing developers to build a custom Playwright stack or a separate automation service, browser capabilities become native tools the AI agent can call whenever a task demands interaction.
What is Browser MCP?
Browser MCP is a specialized MCP server capability that grants AI agents direct control over a headless or managed browser.
It empowers agents to perform actions such as:
- opening URLs;
- clicking elements;
- filling forms;
- typing text and pressing keys;
- reading rendered page content;
- executing JavaScript in the page context;
- scrolling dynamically loaded pages;
- capturing visual screenshots;
- saving pages as PDFs;
- maintaining and reusing authenticated browser sessions.
A typical workflow looks like this:
User task
↓
AI agent
↓
Browser MCP
↓
Managed browser
↓
Website
↓
Clicks / forms / navigation
↓
Result
↓
AI agent
This interaction layer is crucial whenever a website demands active participation rather than simple data extraction.
How Browser MCP works
The interaction flow is highly autonomous:
- The user assigns a browser-based task to the AI agent.
- The agent connects to or launches a managed browser session via its MCP tools.
- The browser navigates to the requested URL.
- The agent reads the rendered page and decides its next move.
- It then clicks, types, navigates, executes JavaScript, or extracts data as needed.
- The system gracefully handles proxies, anti-bot mechanisms, or CAPTCHAs if they appear.
- The final result is returned to the agent.
For example, an authentication workflow might look like:
Open a website
↓
Find the login form
↓
Fill email and password
↓
Handle captcha if required
↓
Submit
↓
Save authenticated session
↓
Reuse it later
This flexibility allows agents to accomplish complex, multi-step tasks that a single HTTP request could never handle.
Browser automation for AI agents
The key difference between an HTTP scraper and Browser MCP is the capacity for interaction.
A standard scraper simply pulls data:
URL
↓
HTML
↓
Data
Browser MCP executes actions:
URL
↓
Open browser
↓
Inspect page
↓
Click
↓
Type
↓
Navigate
↓
Read result
This interactive capability makes Browser MCP perfect for modern web apps, JavaScript-heavy sites, dashboards behind logins, and any workflow demanding a real browser context.
Navigate websites
The foundational navigation tool is:
browser_navigate
It opens a URL inside the managed browser. From there, the agent can inspect the layout and decide how to proceed. A basic sequence often looks like:
browser_navigate
↓
browser_snapshot
↓
browser_click
↓
browser_get_text
The agent doesn't need to be hardcoded with a rigid sequence of steps. It can organically inspect the browser's state after each action and determine what to do next.
Click and interact with elements
Browser MCP provides a suite of tools for engaging with page elements:
browser_click
browser_fill
browser_type
browser_press_key
These allow the agent to click buttons, open dropdowns, enter search queries, complete logins, submit forms, and navigate multi-step flows. Because the session is stateful, the browser context is perfectly preserved between actions:
browser_navigate
↓
browser_fill
↓
browser_click
↓
browser_snapshot
Fill forms
Form automation is a core use case for Browser MCP. The agent can seamlessly:
Open form
↓
Find fields
↓
Enter values
↓
Select options
↓
Submit
↓
Read result
This workflow applies to search bars, account logins, internal dashboards, filter menus, configuration pages, and complex multi-step wizards. Most of these tasks are easily handled by the standard browser tool group, while niche interactions are available via browser_full.
Read page content
Automation requires the agent to understand the current state of the page. Browser MCP provides several distinct ways to read content.
browser_snapshot
This returns an accessibility-oriented, structural representation of the page. It's incredibly useful because it highlights interactive elements without overwhelming the LLM with the entire raw HTML DOM.
browser_get_text
This simply returns the clean, human-readable text currently visible on the page.
browser_get_html
When the agent needs to parse underlying markup or pass it to another tool, this returns the raw HTML of the current state.
A typical read-and-react sequence:
browser_navigate
↓
browser_snapshot
↓
Find relevant element
↓
browser_click
↓
browser_get_text
Execute JavaScript
When standard clicks and typing aren't enough, the browser_full group provides the browser_evaluate tool.
browser_evaluate
This allows the agent to execute custom JavaScript directly within the page context. It's highly effective for reading hidden application states, triggering internal API calls, extracting specific variables, or manipulating page functionality. Generally, JavaScript execution should be a fallback when native browser actions fall short.
Work with dynamic websites
Modern websites often load content dynamically after the initial request via React, infinite scrolling, or client-side APIs. An HTTP scraper will usually just see an empty loading shell.
Because Browser MCP runs a real, managed browser, it waits for JavaScript to execute and render the DOM before the agent attempts to read or interact with the page. This makes it indispensable for modern e-commerce sites, interactive dashboards, and single-page applications.
Persistent browser sessions
Logging in repeatedly during every task is both inefficient and likely to trigger security blocks:
Run 1: Login → Work → End
Run 2: Login → Work → End
Run 3: Login → Work → End
Browser MCP supports session persistence, allowing the agent to authenticate once and reuse that state across future tasks:
Run 1: Login → Save session
Run 2: Load session → Already authenticated
Run 3: Load session → Already authenticated
The core session management tools are:
browser_save_session
browser_load_session
browser_list_sessions
This drastically reduces friction for recurring tasks, dashboard monitoring, and operations on authenticated sites.
Save and restore browser sessions
Once authentication is complete, the agent can freeze the browser state:
browser_save_session
When a new task requires that login, the agent can instantly restore it:
browser_load_session
This stored state includes authentication cookies, ensuring the website recognizes the agent as a logged-in user. You can inspect available sessions using browser_list_sessions. This is a massive time-saver for logins that involve complex multi-factor steps or CAPTCHA challenges.
Browser profiles and authentication
By leveraging persistent state, an authenticated workflow becomes incredibly streamlined:
First run
─────────
Open login page
↓
Enter credentials
↓
Handle captcha
↓
Login
↓
Save session
Later run
─────────
Load session
↓
Open authenticated page
↓
Continue task
(Note: If you are building workflows that directly import cookies or manage Browser API profiles programmatically, the standalone Browser API offers a separate profile mechanism. Browser MCP and Browser API solve similar problems but expose them at different abstraction levels.)
Unlocker with Browser MCP
Browser automation frequently hits roadblocks when it encounters CAPTCHAs during login or navigation. The 2Captcha MCP server integrates built-in tools for detecting and solving these challenges:
detect_captcha
solve_captcha
solve_captcha_on_page
The workflow natively handles interruptions:
Open page
↓
Captcha appears
↓
Detect WAF
↓
Unlocker
↓
Apply solution
↓
Continue browser workflow
Because these CAPTCHA tools are separated from the core browser group, they can also be used independently.
Use Browser MCP with your own browser
You aren't forced to use the managed Browser MCP browser. If your existing stack already uses Playwright, browser-use, or a custom Chrome extension, you can maintain your own environment and exclusively call the 2Captcha MCP tools for CAPTCHA resolution:
Your browser
↓
Captcha detected
↓
Page HTML
↓
detect_captcha
↓
solve_captcha_on_page
↓
Solution
↓
Apply to your browser
However, if you prefer the MCP server to manage the entire lifecycle of the browser, simply enable the full Browser MCP tool groups.
Proxy support
Browser sessions can route traffic through proxies when a task dictates it. This is vital for workflows that depend on geographic localization, bypassing IP restrictions, maintaining network reputation during heavy use, or researching regional marketplaces. Proxies, anti-bot handling, and unlocker can all operate simultaneously within a single workflow. See 2Captcha Browser MCP for configuration options.
Browser MCP tool groups
Browser automation is strictly opt-in, organized into two distinct groups based on the level of control required:
browser
browser_full
browser tool group
This group covers the essentials for standard workflows:
- Navigation/Interaction:
browser_navigate,browser_click,browser_fill,browser_type,browser_press_key - Reading pages:
browser_snapshot,browser_get_text,browser_get_html - Sessions:
browser_save_session,browser_load_session,browser_list_sessions
browser_full tool group
For Playwright-style granular control, browser_full adds advanced actions:
browser_go_back
browser_go_forward
browser_reload
browser_scroll
browser_select_option
browser_hover
browser_drag
browser_console_messages
browser_evaluate
browser_snapshot_items
browser_screenshot
browser_save_as_pdf
The latest tool specs are maintained in the official 2Captcha MCP GitHub repository.
Capture screenshots and PDFs
The browser_full group supports visual capture via browser_screenshot. This is highly useful for debugging, verifying state changes, or inspecting visual UI bugs.
Additionally, browser_save_as_pdf allows the agent to generate and preserve formatted PDF documents like receipts, financial reports, or official documentation.
Browser MCP API
The hosted endpoint is:
https://mcp.2captcha.com/mcp
Authenticate using your 2Captcha API key as a bearer token:
Authorization: Bearer YOUR_API_TOKEN
For current connection guidelines, visit the 2Captcha Browser MCP page.
How to connect Browser MCP
Hosted Browser MCP
Connect directly without running local Node.js processes:
URL: https://mcp.2captcha.com/mcp
Header: Authorization: Bearer YOUR_API_TOKEN
Local Browser MCP
Launch the official package locally using npx @2captcha/mcp.
To load the standard browser tools in your client:
{
"mcpServers": {
"2captcha": {
"command": "npx",
"args": ["@2captcha/mcp"],
"env": {
"API_TOKEN": "YOUR_API_TOKEN",
"GROUPS": "browser"
}
}
}
}
Change "GROUPS": "browser" to "GROUPS": "browser_full" if your agent requires the advanced toolset. Detailed examples are available in the 2Captcha MCP GitHub repository.
Browser MCP with Claude Code, Cursor, and Codex
Claude Code can connect to the hosted endpoint:
claude mcp add --transport http 2captcha https://mcp.2captcha.com/mcp \
--header "Authorization: Bearer YOUR_API_TOKEN"
Or run locally:
claude mcp add 2captcha -e API_TOKEN=YOUR_API_TOKEN -e GROUPS=browser -- npx @2captcha/mcp
Cursor directly supports the local setup via its MCP config panel, seamlessly adding browser navigation to its coding capabilities.
Codex CLI functions similarly:
codex mcp add 2captcha --env API_TOKEN=YOUR_API_TOKEN --env GROUPS=browser -- npx @2captcha/mcp
Browser MCP GitHub repository
The official implementation, technical docs, tool groups, and configuration examples for Claude, Cursor, VS Code, and Codex are all public:
Run the official local package with:
npx @2captcha/mcp
Browser MCP use cases
Authenticated website automation
Log in once, save the session, and seamlessly resume authenticated workflows on subsequent runs.
Dynamic website navigation
Bypass the limitations of HTTP scraping by rendering JavaScript-heavy pages before extracting data.
Form and Dashboard automation
Navigate deeply nested dashboards, fill out complex forms, select filters, and respond dynamically to website feedback.
Data extraction after interaction
Execute a search, apply visual filters, click through pagination, and open specific items before extracting the final structured data.
Protected website workflows
Merge proxy routing, unlocker, and session persistence to navigate sites wrapped in aggressive anti-bot protections.
Browser MCP vs MCP Scraper
| MCP Scraper | Browser MCP |
|---|---|
| Retrieves page content | Controls a browser |
| Starts with a URL | Starts with a browser task |
| Optimized for extraction | Optimized for interaction |
| No step-by-step browser control required | Agent controls browser actions |
| Better for data retrieval | Better for multi-step workflows |
Use MCP Scraper for pure data extraction. Use Browser MCP when you need to actively interact with the page.
Browser MCP vs MCP Web Search
MCP Web Search is for discovery:
Query → Search results → URLs
Browser MCP is for interaction:
URL → Browser → Clicks / typing / navigation → Result
For page discovery, rely on 2Captcha Web Search MCP.
Browser MCP vs Web MCP
Web MCP is the umbrella toolkit that encompasses Search, Scraping, Extraction, Browser automation, and unlocker. Browser MCP is simply the dedicated interactive component of that broader suite.
Browser MCP vs Browser API
Browser MCP exposes browser functionality as high-level tools for an AI agent to call autonomously.
Browser API provides programmatic control over a managed browser directly for your own code to execute.
Browser MCP security and saved sessions
Because saved sessions contain active authentication cookies, they effectively act as user credentials. Treat them with strict security:
- Only save sessions when persistence is strictly required.
- Never expose session data to unrelated systems.
- Routinely review and delete obsolete stored sessions.
- Rotate credentials if you suspect a session has been compromised.
FAQ
What is Browser MCP?
It provides AI agents with fully managed browser automation capabilities through the Model Context Protocol, enabling them to navigate, interact with, and read websites dynamically.
What can Browser MCP automate?
It automates navigation, clicking, typing, form filling, JS execution, screenshots, PDF rendering, and session management.
Which Browser MCP tools are available?
The browser group offers core navigation and session tools, while browser_full adds advanced Playwright-style interactions.
Can Browser MCP work with dynamic websites?
Yes. It runs a real browser that fully executes JavaScript and renders dynamic content.
Can Browser MCP keep login sessions?
Yes. Use browser_save_session and browser_load_session to reuse authenticated states across multiple tasks.
Can Browser MCP handle captcha?
Yes. The 2Captcha MCP server includes integrated tools for detecting and solving CAPTCHAs during browser workflows.
Does Browser MCP support proxies?
Yes, it supports proxy configurations for workflows needing specific geographic routing or IP rotation.
Can I use Browser MCP with Claude, Cursor, or Codex?
Yes. All major MCP-compatible clients support connecting to either the hosted endpoint or the local @2captcha/mcp package.
Can I use my own Playwright browser?
Yes. You can manage your own Playwright instance and solely invoke the 2Captcha MCP tools when you encounter a CAPTCHA.
What is the Browser MCP API endpoint?
The hosted endpoint is https://mcp.2captcha.com/mcp, authenticated via the Authorization: Bearer YOUR_API_TOKEN header.
Where is the Browser MCP documentation?
Product overviews are at 2Captcha Browser MCP. Full technical docs and source code live in the official 2Captcha MCP GitHub repository.
What is the difference between Browser MCP and Browser API?
Browser MCP is built for autonomous AI agents deciding on actions; Browser API is for developers writing deterministic programmatic automation.