AI MODELS / BROWSER AUTOMATION
AI models for
browser automation.
Research sites, collect data, and fill out forms with rtrvr. Compare model speed, cost, and results to choose one for your task.
Compare models at a glance
Benchmark score, speed, cost, and context for up to five models.
Index and API speed: Artificial Analysis, 23 September 2026 (Artificial Analysis Intelligence Index v4.3.2, maximum reasoning effort). Cost uses the everyday sample task. Models without an AA score are in the data table.
Live: how models run as browser agents
Real traffic, not a benchmark. Speed, reliability, and task outcomes for every model and provider.
- Most used69%of callsGPT-6 Luna
- Fastest first token1.0 smedianDeepSeek Flash
- Fastest output223tokens/secDeepSeek Flash
- Fewest failed responses0.3%of callsGPT-5.6 Luna
Share of agent calls
Every hour, across all browser-agent traffic- GPT-6 Luna 69%
- GLM 5.3 Flash 16%
- MiMo V2.6 Flash 4%
- DeepSeek Flash 4%
- Other 7%
| Model · provider | Shareof calls | First tokenmedian | First tokenp90 | Speedtokens/sec | Responsemedian | Errorsgeneration | Cache hitof input | Answeredact runs | Doneplan runs | Reviewedplan runs | Stuckact runs |
|---|---|---|---|---|---|---|---|---|---|---|---|
GPT-6 Luna | 69% | 1.6 s | 4.3 s | 89 | 2.5 s | 0.5% | 84% | 98% | 81% | 59% | 3% |
OpenAI | 68% | 1.2 s | 2.2 s | 82 | 1.9 s | 0.1% | 85% | 99% | 81% | 61% | 3% |
Amazon Bedrock | 31% | 2.7 s | 7.0 s | 93 | 5.0 s | 1.5% | 85% | — | 81% | 59% | — |
ChatGPTuser key | <1% | 1.7 s | 3.0 s | 34 | 8.2 s | 0.0% | 28% | 90% | 90% | 90% | 0% |
GLM 5.3 Flash | 16% | 2.0 s | 4.1 s | 85 | 4.0 s | 0.5% | 75% | 84% | 88% | 34% | 2% |
Azure | 93% | 2.0 s | 3.9 s | 84 | 4.0 s | 0.4% | 77% | 85% | 88% | 35% | 2% |
Particle | 5% | 1.5 s | 3.2 s | 202 | 2.5 s | 0.0% | 59% | — | — | — | — |
Z.ai | 2% | 6.0 s | 13 s | 48 | 8.5 s | 0.0% | 33% | — | 100% | 40% | — |
MiMo V2.6 Flash | 4% | 4.4 s | 16 s | 45 | 18 s | 2.1% | 73% | 88% | 73% | 0% | 0% |
Xiaomi | 100% | 4.4 s | 16 s | 45 | 18 s | 1.7% | 73% | 88% | 73% | 0% | 0% |
DeepSeek Flash | 4% | 1.0 s | 1.8 s | 223 | 3.7 s | 1.5% | 70% | 88% | 74% | 5% | 0% |
Particle | 98% | 1.0 s | 1.8 s | 223 | 3.7 s | 0.7% | 70% | 88% | 74% | 5% | 0% |
GPT-5.6 Luna | 3% | 2.2 s | 3.9 s | 50 | 4.4 s | 0.3% | 64% | 86% | 54% | 8% | 36% |
GPT-5.6 Terravia ChatGPTuser key | <1% | 2.0 s | 4.6 s | 42 | 7.1 s | 1.2% | 41% | — | — | — | — |
Gemini Flash Litevia Googleuser key | <1% | — | — | 93 | 3.9 s | 0.0% | 21% | — | — | — | — |
Gemini Flashvia Google | <1% | — | — | 236 | 3.5 s | 6.7% | 13% | — | — | — | — |
Medians from real browser tasks in the extension, Cloud, and API. Errors count failed model responses, not invalid keys or exhausted quotas. A row appears after five people have used it, so provider shares can total under 100%; a dash means too little data. Answered and Done are what the agent reported. Data sources
Performance by provider
Every provider serving a model, hour by hour.
Output speed
Output tokens per second, median- OpenAI82 tok/s
- Amazon Bedrock93 tok/s
- ChatGPTuser key34 tok/s
First token
Time to the first output token, median- OpenAI1.2 s
- Amazon Bedrock2.7 s
- ChatGPTuser key1.7 s
Round time
Median time for one model response- OpenAI1.9 s
- Amazon Bedrock5.0 s
- ChatGPTuser key8.2 s
Failed responses
Calls that failed to generate- OpenAI0.1%
- Amazon Bedrock1.5%
- ChatGPTuser key0.0%
Every model, side by side
Prices, sample task costs, benchmarks, and our historical measurements.
AA benchmarks: 23 September 2026. Historical rtrvr data: 31 August 2026. Open a row for notes. How we measure
| Task cost · credits | Price per 1M tokens · USD | Artificial Analysis | Historical rtrvr sample | Inputs | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
DeepSeek FlashFree ModeCollecting data across pages; reads images | 0.08 | 0.90 | 3.0 | $0.285 | $0.006 | $1.14 | No separate cache-write rate listed | 39.5 | 232 | 12.9 | 85.3 | 72% | 77 | 1M | Yes |
DeepSeek ProTasks with longer reasoning | 0.25 | 3.2 | 11 | $1.056 | $0.035 | $3.168 | No separate cache-write rate listed | 36.0 | 67 | No comparable measurement in this sample | No comparable measurement in this sample | 60% | 116 | 128K | No |
Gemini Flash LiteFree ModeQuick lookups and summaries | 0.12 | 1.2 | 4.4 | $0.30 | $0.03 | $2.50 | No separate cache-write rate listed | 22.2 | 367 | No comparable measurement in this sample | No comparable measurement in this sample | 42% | 209 | 1M | Yes |
Gemini FlashPages with images and PDFs | 0.22 | 2.6 | 9.4 | $0.75 | $0.075 | $3.75 | No separate cache-write rate listed | Not rated | 315 | No comparable measurement in this sample | No comparable measurement in this sample | 34% | No comparable measurement in this sample | 1M | Yes |
GPT-6 LunaFree ModeQuick questions, tool calls, and images | 0.03 | 0.35 | 1.3 | $0.10 | $0.01 | $0.50 | $0.125 | 37.3 | 162 | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 1M | Yes |
GPT-6 SolComplex sites and extensive tool use | 0.60 | 7.0 | 25 | $2.00 | $0.20 | $10.00 | $2.50 | 47.5 | 131 | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 1M | Yes |
Inkling SmallFree ModeLow-cost text-based tasks | 0.08 | 1.1 | 4.1 | $0.30 | $0.06 | $1.20 | No separate cache-write rate listed | 27.8 | 199 | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 256K | No |
InklingText tasks with longer reasoning | 0.27 | 3.5 | 13 | $1.00 | $0.17 | $4.05 | No separate cache-write rate listed | 25.0 | 114 | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 256K | No |
MiMo V2.6 FlashFree ModeLong text with repeated context; processed in China | 0.03 | 0.40 | 1.3 | $0.14 | $0.003 | $0.28 | No separate cache-write rate listed | Not rated | Not measured | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 1M | Yes |
MiMo V2.6 ProLong documents and deeper reasoning; processed in China | 0.09 | 1.2 | 3.9 | $0.435 | $0.004 | $0.87 | No separate cache-write rate listed | Not rated | Not measured | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 1M | Yes |
GLM 5.3 FlashFree ModeCode, tool calls, and screenshots | 0.02 | 0.32 | 1.2 | $0.09 | $0.018 | $0.30 | No separate cache-write rate listed | 41.8 | 64 | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 1M | Yes |
GLM 5.3Complex text-based tasks | 0.24 | 3.4 | 13 | $0.98 | $0.182 | $3.08 | No separate cache-write rate listed | 44.8 | 61 | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | No comparable measurement in this sample | 1M | No |
1 credit = $0.01. Scores use Artificial Analysis Intelligence Index v4.3.2 at maximum reasoning effort; rtrvr uses a lower effort by default. A dash means not rated or not measured.
Which model should you try?
Free Mode models, picked by job. Free Mode includes sponsored cards and daily fair-use limits. See plans and limits
Answer questions about a page
Use GPT-6 Luna for questions and tool calls. Try Gemini Flash Lite when output speed matters most. Compare current response times in the live results above.
Collect data across many pages
Try DeepSeek Flash for collecting a dataset, filling a spreadsheet, or working through a list. It supports screenshots. Check live round times if you need a quick response at every step.
Read screenshots, PDFs, and charts
These models accept images. Choose one when the answer depends on a chart, screenshot, or scanned document.
Read long pages and documents
Both support a one-million-token context window and accept images. MiMo V2.6 Flash offers a larger discount on cached input; Xiaomi processes its requests in China.
Spend fewer credits on model tokens
GLM 5.3 Flash has the lowest estimated model cost for the everyday sample workload. It also accepts images. Browser and proxy usage are billed separately on paid runs.
Where the numbers come from
Live results show how models run on rtrvr. Benchmarks and sample costs help you compare them.
Real browser tasks
We update these results hourly. Each model runs a different mix of tasks. The agent reports whether it finished; we do not independently check each result. We omit small samples.
Capability and output speed
Artificial Analysis supplies the index and first-party API speeds, recorded 23 September 2026. The index is a general benchmark, not a browser-task success rate.
Three sample tasks
These examples use token counts from real tasks through 31 August 2026. Estimates cover model tokens. Browser, proxy, and other usage can add to the total. Displayed rates use GMI Cloud for GLM and DeepSeek; OpenRouter for MiMo, which Xiaomi serves from China; OpenAI for GPT-6; Thinking Machines for Inkling; and Google for Gemini. With China routing on in the model picker, GLM runs first on Z.ai in China at half the displayed rate.
Quick lookup
Read a page and answer a question
- Rounds
- 1
- Input
- 1.5K tokens
- Cached
- 0%
- Output
- 300 tokens
Everyday task
Collect data or take a few actions
- Rounds
- 1–2
- Input
- 50K tokens
- Cached
- 50%
- Output
- 1.5K tokens
Long agent run
Complete a task across several pages
- Rounds
- 5+
- Input
- 250K tokens
- Cached
- 70%
- Output
- 6.5K tokens
Try a model on your task.
Choose a model in Cloud. Run the same task with another model to compare results.