browser-use Local Web Agents — VRAM Tiers for Reliable Browser Automation

By Billy G.R. · 2 October 2026

Key facts

Checked
2026-10-02
Author
Billy G.R.

Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.

Citeable facts

Claim. browser-use is the open-source library at github.com/browser-use/browser-use. The repository description is “Agents that use the browser.”

Method. GitHub page read 2 October 2026.

As of.

Claim. Text-only browsing can use a 7–14B Q4 model (6.2–10.4 GB). Screenshot agents should be planned at 24 GB for a 27–30B multimodal model.

Method. Site formula. Qwen3.8-27B is 27B. Muse Glimmer is 30B and includes a perception encoder, per Meta’s August 2026 post.

As of.

Claim. I did not re-confirm the old OpenHands path docs.openhands.dev/sdk/guides/agent-browser-use after the docs split.

Method. OpenHands docs index now leads with Agent Canvas and the SDK. Link the index, not a path I did not re-fetch.

As of.

Methodology — parameter math, quantization bytes, and the source list.

browser-use is an open-source library that lets a language model drive a browser. It is also the kind of tool OpenHands and other agent scaffolds call when there is no API. There is no separate “browser-use appliance” to buy. You install the library, you run Chromium through Playwright, and you point it at a model. The commercial cousin people compare it with is Perplexity Comet, which brings a search index this library does not have.

Reliability is mostly “did the model fit, and can it see”

A browser agent fails in two boring ways. The page changed, so the action is wrong. Or the model was swapped onto the CPU halfway through, so the loop takes minutes per click and you kill it. I can help with the second failure. I cannot publish a success percentage for the first. Sites change. A local model does not get a special exemption.

Decide up front whether the model reads extracted text or a screenshot. Text is cheaper and blind to layout. Screenshots need a multimodal model and a bigger card. Muse Glimmer’s research post is explicit that a perception encoder sits on top of the language-model weights. Do not budget Glimmer as “17 GB and done” if you are feeding it images. Meta’s envelope for language model plus KV plus perception plus the drafter is 24–32 GB.

Tiers

Q4_K_M formula: params × 0.6 + 2. Chromium’s RAM is extra. On unified memory it comes out of the same pool. The calculator will not show the browser process.

What the model sees Weights Card
DOM text, simple clicks 7B = 6.2 GB, 14B = 10.4 GB 12 GB can hold the 14B weights. It is not roomy once Chromium is open on a shared-memory machine.
Screenshots 27B = 18.2 GB, 30B formula = 20 GB 24 GB discrete, or 32–36 GB unified. 16 GB is the wrong purchase.
Long task, many steps Same weights, larger KV cache Add memory for the trace of the run, not a second model. 36 GB is kinder than 24 GB.

An editorial 1–4 GB on top of the formula, for the Playwright process and the agent scaffold, is a planning fudge. It is not an NVIDIA spec. The agent formula guide says the same thing in a table.

A card that matches screenshots

If screenshot automation is the actual job, buy 24 GB, not 16 GB. The catalog RTX 4090 is ASIN B0BG94PS2F. A 12 GB RTX 5070 (ASUS TUF, B0DS6S98ZF) is a text-only 7–14B machine; the listing showed $937.39 on 2 October 2026 and it does not grow a vision model down to 12 GB. On Mac, the Studio 36 GB (B0HGKSQMX6) is the unified-memory version of “screenshots plus cache.”

RTX 4090 24 GB on Amazon

ASIN B0BG94PS2F. Amazon Associates link. The live price is on that page, not locked in here.

ASUS TUF RTX 5070 12 GB on Amazon

ASIN B0DS6S98ZF. Amazon Associates link. The live price is on that page, not locked in here.

Related guides

Sources

A correction or a hardware question goes to the contact page.