firefox-use
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| FIREFOX_BIN | No | Path to the Firefox executable. | |
| FIREFOX_USE_PORT | No | Port for the WebDriver BiDi server. | |
| FIREFOX_USE_PROFILE | No | Path to a Firefox profile directory. | |
| FIREFOX_USE_HEADLESS | No | Set to '1' to run Firefox headless. | |
| FIREFOX_USE_DOWNLOADS | No | Directory for downloads. | |
| FIREFOX_USE_AUTO_RESTART | No | Set to '0' to disable auto-restart. | |
| FIREFOX_USE_UPLOAD_ROOTS | No | Root directories allowed for uploads. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| navigateA | Send a tab to an address, or walk that tab through the history it has already built. Only url is always required. For an ordinary address you may leave tabId out when navigate is called on its own: the session's tab group is opened or reused for you (the work tabs_context_mcp{createIfEmpty:true} does), the group's first tab is driven, and the refreshed listing is printed under the result so the ids are in front of you for the next call. Two cases have no fallback and fail without a tabId: url set to "back" or "forward", and any navigate running inside browser_batch, where a guessed tab would be the one that existed before the batch opened its own. Name the tab yourself whenever the group holds several pages and the others are mid-task. The call returns after loading settles, which retires every ref_N an earlier read_page handed out - read the page again before using refs. In Firefox the page opens in the automation profile this server launched, apart from your everyday windows and with no per-site approval step; a load parked behind an alert or a beforeunload prompt is cleared with firefox_dialog instead of ending the session. |
| resize_windowA | Change the size of the Firefox window a tab lives in, so a layout can be checked at phone, tablet or desktop width and so screenshots taken afterwards come out at that geometry. width, height and tabId are all required; there is no implicit tab here. The real window is resized when Firefox allows that; headless runs and older builds refuse that, and the tab's viewport is stretched instead, which puts the page at the same size. The reply names whichever path ran and quotes the innerWidth x innerHeight the page ended up reporting - expect it to fall short of the numbers you asked for, since toolbars come out of the window box and the browser clamps sizes it thinks are too small. Only the tab you name is affected; the rest of the group keeps rendering as it was. |
| javascript_toolA | Evaluate JavaScript inside a page and hand back what its final expression produced. All three fields are required: action pinned to javascript_exec, text carrying the source, and tabId naming the page. The snippet runs in the page's own world, so the document, window, and whatever globals the site defined are all within reach, and anything it changes is really changed. It behaves like a console rather than a function body: top-level await is allowed, and the value of the last expression comes back by itself, so finish with the expression instead of a return statement. Results are stringified (objects as indented JSON) and cut off at 10000 characters, with a line saying how long the full text was. A throw from the page fails the call with its message and the frame that raised it. Two things script cannot do here: fill a file input, which needs file_upload or upload_image, and get past its own alert or confirm - that parks the page until firefox_dialog answers it, which in Firefox is recoverable rather than the end of the session. |
| form_inputA | Fill one form control that read_page has already tagged with a ref, without aiming a click at it. ref, value and tabId are all required. Text boxes, textareas, number and date fields, contenteditable regions, checkboxes, radios and dropdowns are covered; a file input is not, and goes through file_upload or upload_image, which read only from the directories this server is allowed to open. The write goes through the element's native setter and is followed by input and change events, while checkboxes and radios are toggled with a real click, so React and its relatives notice the change. A dropdown is matched against option values first, then against the labels shown to the user, ignoring case on a last pass; when nothing matches, the error lists the options that exist. Disabled controls are turned down. A ref lives only as long as the node behind it, so after a navigation, a reload, or a re-render that swaps the element out, run read_page again or this fails with an unknown-ref error. |
| tabs_context_mcpA | List the tabs this session owns - numeric id, current address, page title, and which one holds focus. Run it before any other browser tool, because they all address a page by an id from this list, and run it again after tabs open or close. The group only ever holds tabs this server opened in the Firefox profile it launched, so the windows you browse in yourself are neither listed nor driven, and no site has to grant automation permission first. When no group exists yet the reply says so and points at createIfEmpty. Give a new conversation a tab of its own with tabs_create_mcp rather than taking over one that is already busy, unless the user asked you to work in the tab that is open. |
| tabs_create_mcpA | Open one more blank tab (about:blank) in the session tab group and report its id together with what the group now holds. It takes no arguments. The tab joins whichever Firefox window the group already occupies, or brings one of its own when the group is empty. Call tabs_context_mcp once first, so you know the group you are adding to, then put the id you were just given on a navigate to load something into it. Whatever you open is yours to clear away: close it with tabs_close_mcp the moment it stops being useful, and leave none of your own tabs behind when the task ends - keep one only when the user wants to look at it. |
| tabs_close_mcpA | Shut one tab out of the session tab group, addressed by id - the way you tidy up once a page has served its purpose. tabId is required, and has to be an id tabs_context_mcp reported: a tab from one of your own windows, or one that is already gone, is refused rather than quietly ignored. Ids are handed out in sequence and never reused, so a closed one stays invalid for the rest of the session. Shutting the last remaining tab empties the group and Firefox drops it; a later tabs_context_mcp that passes createIfEmpty, or any tabs_create_mcp, builds a fresh one. |
| read_pageA | Dump a loaded page as an indented accessibility tree: one line per element carrying its role, accessible name, state, its viewport box when it has one, and a ref_N handle, which is what computer, form_input and the other element tools take as a target. Everything is listed by default, including elements scrolled out of view or hidden by CSS; pass filter "interactive" to keep only the controls. The walk stops 15 DOM levels below the body unless depth raises it, gives up after 5000 elements, and the rendered text is trimmed to 50000 characters (max_chars) on a line break so that no handle is cut in half - whenever any of those three fires, a bracketed note at the end names the limit that hit, how much of the whole you got and how to reach the rest. Handles are minted per document: after a navigation or a reload the old ones stop resolving, so read the page again instead of reusing them. tabId is required (tabs_context_mcp or firefox_status lists the ids this session owns), and a step inside browser_batch has to carry it too. Firefox specifics: the browser runs in its own automation profile, so nothing asks you to approve a site, but an open alert or confirm freezes the page until firefox_dialog answers it. |
| findA | Look up elements by describing them in ordinary language instead of by selector: "the search box at the top", "whatever button adds this to a cart", "the heading that mentions pricing". Scoring runs over the same tree read_page builds - accessible names, an element's own text, and attributes such as id, name, placeholder, title and href - with extra weight for a handful of intents (search, login, cart, menu, close, submit, dropdown, checkbox, heading). Results come back ranked, best first, each with the ref_N handle the action tools need, and at most 20 are printed; when more than that matched, the output tells you to describe the target more tightly. The search reaches 30 levels deep and the first 5000 elements walked, and it says so when it stopped short - so an empty result on a huge page means widen the net with read_page, not that the element is absent. query and tabId are both required, and handles expire with the document, so run find again after any navigation; a browser_batch step must state the id itself. As everywhere on this server the automation profile skips any per-site approval step, while an unanswered alert or confirm holds the page until firefox_dialog clears it. |
| get_page_textA | Read a page as prose: plain text, no markup, no element handles. The extractor takes the article bodies when the page has real ones, otherwise , otherwise the block with the best text-to-markup ratio, falling back to the body, and throws away navigation, banners, footers, sidebars, toolbars, select menus, iframes and anything marked hidden or aria-hidden, so a long post arrives as paragraphs rather than as the furniture around them. Reach for it when the goal is to read or quote; reach for read_page or find when you need something clickable, since nothing here carries a ref. Output stops at 200000 characters, cut on a line break, and the closing note reports how long the page actually was - only a pathological page ever hits that. tabId is required, must name a tab this session owns, and has to be spelled out on a browser_batch step. A modal alert or confirm blocks the read until firefox_dialog answers it; no site approval is asked for, since the browser runs in its own profile. |
| computerA | Drive one Firefox tab with synthetic pointer, wheel and keyboard input, and capture what it looks like. Each call performs exactly one
|
| read_console_messagesA | Hand back the console output firefox-use has buffered for one tab: log, info, warn and error entries, plus uncaught page exceptions with their stack frames. Every line arrives as an ISO timestamp, the level in brackets and the text, oldest first, and an error also shows the frame it was thrown from. Capture is lazy - the buffer only starts filling the first time you call this tool or read_network_requests in a session, so anything printed before that was never seen; reload and read again. A tab keeps its 1000 most recent entries, and the whole log is dropped when that tab moves to a different origin (a new scheme, host or port each count), so read before you send the page somewhere else; the reply says so when either of those cost you entries. tabId is the only required argument and has to name a tab this session owns - tabs_context_mcp lists them. Send a pattern nearly every time: on a talkative app an unfiltered read buries the two lines you were after. |
| read_network_requestsA | List the HTTP traffic recorded for one tab - documents, scripts, stylesheets, images, XHR and fetch alike - as one padded row per request, oldest first, giving method, status, resource type, elapsed time, bytes transferred and URL, flagged when the request failed or is still in flight. Use it to confirm an API call actually fired, see what came back, or find the asset costing you seconds. Traffic to any host is kept, not only the page's own origin, but a tab's log is emptied as soon as that tab moves to a different origin (scheme, host or port), and only its 1000 most recent requests survive. Recording begins with your first call to this tool or read_console_messages, so a request made earlier than that does not exist here; reload and read again. tabId is required and has to be a tab in this session's group - tabs_context_mcp will give you one. |
| file_uploadA | Puts one or more files from this machine into an on the page. Never try to click the input or the button that fronts it: Firefox answers a click with the operating system's own picker window, which sits outside the page where nothing here can reach it. Locate the input with read_page or find instead, then call this tool with the ref it was listed under. All three arguments - paths, ref and tabId - are required. Reading a file is gated: a path is accepted only when it resolves inside an upload root, meaning this session's download folder, the directory the server process was started in, or an absolute path named in the FIREFOX_USE_UPLOAD_ROOTS environment variable (separate several with the platform path separator). Symlinks are followed before that check, so a link parked in a root but aimed outside one is turned down as well, and the refusal lists the roots that would have worked. Chrome instead scopes uploads to whatever the user shared with the conversation; this server has no such notion. One call may carry less than 10 MB in total across every file named. Afterwards the input is read back and the result states how many files it now holds, which is how you catch an input with no multiple attribute keeping just one of several. |
| upload_imageA | Feeds a picture the session is already holding - a screenshot or zoom taken with the computer tool, which prints an imageId such as img_3 - into a page, so no file has to travel through you. There are two ways to aim it and you must pick exactly one: ref, an element from read_page or find, which is how you reach a file input the page keeps hidden behind styling; or coordinate, an [x, y] point in the viewport, for editors that take a dropped file and expose no input at all. Giving both is an error and so is giving neither; imageId and tabId are always needed. Down the ref path the image is written to a scratch file in this session's download folder and handed to the input, and should that element turn out not to be a file input, the tool scrolls to it and drops the image on it instead. Down the coordinate path whatever sits under the point gets a synthesized dragenter, dragover and drop, plus a paste event as a fallback when none of the three was taken up. Read the result line: it names which of drop, paste or nothing the page actually handled, so a target that quietly ignored the image cannot pass for success. |
| gif_creatorA | Films a tab as a run of screenshots and encodes them into an animated GIF, with the clicks and drags of the run marked on the frames that show what they did. Four actions, and every one of them needs tabId because a recording belongs to a single tab: start_recording opens the loop and grabs a first frame right away, stop_recording ends the loop after one closing frame and keeps everything captured, export encodes the frames into a .gif on disk, clear throws them out. The loop grabs a frame about every 700 ms and halts itself once 120 are in hand; frames are shrunk to at most 800 px wide, and the delay stored in the GIF follows the cadence that was really measured, not the nominal one. Calling export while the loop still runs stops it first, and the frames outlive the export - clear them when you are done, or the next start_recording carries on the same set. Frames live in server memory only: a tab that closes mid-recording ends the loop with a note in the next reply, and any frame whose size differs from the first (a window resize partway through) is dropped at encode time and counted in the result. Where Chrome pushes the finished GIF through the browser download machinery, this server writes it into the session's download folder itself and gives you back the full path. |
| browser_batchA | Push a whole run of browser tool calls through one request rather than spending a turn on each. Every entry is {name, input}, where input is the same object you would hand that tool on its own. Entries run one after another, never in parallel, and the run halts at the first failure: you get the output of everything that finished, the error from the entry that broke, and a count of the entries that never started. The list is checked before anything touches the page, so an unknown tool name, or an attempt to put browser_batch inside itself, rejects the call outright instead of half-applying it. Batch aggressively whenever you can see two or more moves ahead - open the page, click into the search box, type the query, press Return, take a screenshot. Images come back inline among the text outputs, each labelled with the entry that produced it. Two things bite here. First, any coordinate you write in these entries, and any ref you name, come from a screenshot taken BEFORE the call or a read_page that ran before it, since nothing this batch paints is visible to you while you compose it - and once an entry navigates or reloads, every earlier ref and coordinate is stale, so end the batch there and look again. Second, every entry that acts on a page has to state tabId - inside a batch no tab is inferred, so a batch that creates a tab and then drives it must repeat that new id on each later entry. Unlike the Chrome extension, firefox-use runs no per-site permission check between entries; it drives its own automation profile. A modal dialog does still freeze the page, so clear it with firefox_dialog rather than letting the rest of the batch queue up behind it. |
| firefox_statusA | Report whether Firefox is running, which profile, port and viewport it uses, and which tabs are in this session's group. Firefox-only; the Chrome integration has no equivalent. |
| firefox_restartA | Restart Firefox. Use this when the browser stops responding or reports a stale automation session. Firefox-only. |
| firefox_dialogA | Accept or dismiss a native dialog (alert, confirm, prompt, beforeunload) that is blocking the page. Firefox-only: in Chrome a blocking dialog ends the automation session, here it can be cleared. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 20 tools
Most tools are clearly separated by role: tab management, page reading, form input, file handling, network/console logging, and Firefox maintenance. The main overlap is that tabs_context_mcp and firefox_status both report the session's tab list, and read_page, find, and get_page_text have somewhat similar page-reading purposes even though the descriptions distinguish them well.
Although most names use snake_case, there is no consistent pattern across the set. Tab tools use an odd _mcp suffix, several tools are bare or generic nouns like find, computer, and gif_creator, and there is no uniform verb_noun or noun_verb convention across the rest. A new agent cannot reliably guess related tools from their names.
20 tools is on the heavy side and somewhat at the edge where complexity starts to overwhelm. The count reflects a broad browser-automation scope, but some responsibilities are spread across multiple tools rather than being consolidated, so the set feels larger than necessary even though it is not unreasonable.
The tool surface covers a nearly complete browser workflow: tabs, navigation, reading and finding elements, form input, file and image uploads, screenshots, keyboard/pointer control, JS evaluation, console and network logging, GIF capture, and Firefox restart/dialog handling. Missing pieces like cookie/storage management or explicit screenshot download are minor and usually workable with changes or extensions.