so this is the follow up to the last post. if you haven't read that one, the short version is, i built a local mcp that gives amazon q the ability to browse the internet. i called it qbrowser.
this post is about what it actually became.
the numbers
i did a side by side comparison of qbrowser against claude's built-in browser tools. the results were kind of wild to look at.
109 vs 17. that's not a small gap. and it's not just quantity, the categories are different too.
what qbrowser has that claude doesn't
this is the part that actually surprised me when i laid it all out. things qbrowser can do that neither of claude's browser modes can:
- network interception, block, mock and modify any request
- full storage access, cookies, localStorage, sessionStorage, IndexedDB
- human-in-the-loop input modal for things like CAPTCHAs
- session persistence, save and reload full browser state
- auto-retry any tool with configurable backoff
- smart wait, selector, text, network idle, DOM stable
- switch engine, chromium, firefox or webkit
- stealth mode, anti-bot fingerprint spoofing
- iframe interaction, list, click, execute JS inside
- performance metrics and core web vitals
- accessibility audit via axe plus color contrast check
- geolocation spoofing
- mp4 native video recording
- pdf export
- clipboard read and write
- visual snapshot diff, before, diff, after
like, that's a lot. and most of that came from amazon q itself when i handed it the v1 and asked it to amplify the system. it knew what was missing and just built it out.
what claude has that qbrowser doesn't
to be fair, claude's browser isn't useless. it has things i don't:
- dev server preview built in, start, stop, read logs
- zero setup, works immediately inside claude code
- sandboxed, browser actions can't touch your system state
- claude-in-chrome connects to your real chrome, already logged in to everything, real extensions active
- batch multiple actions in one round trip
- native gif recording
the real chrome thing is genuinely useful. if you're already logged into something, you don't have to deal with auth at all. that's a real advantage.
the hosted mvp situation
i hosted the mvp on a server at some point to see how it would run externally. and it worked, kind of. the problem was i couldn't see the actual chromium instance doing its thing live. when it's running locally you can watch it, you can see the browser moving, clicking, navigating in real time. when it's hosted externally and accessed remotely, that visual feedback is gone. you're just reading logs and hoping.
so it ran on its own, which was cool. but the whole point of having a browser agent is being able to watch it work and course correct. without that it felt like flying blind again, which is ironic given the whole premise of this project.
that's something i need to figure out for v3. some kind of live view or stream of what the browser is doing when it's running on a remote server.
where this is going
v2 is solid. 109 tools, runs locally, amazon q can navigate, read, interact, record, audit, spoof, intercept, the whole thing. it's genuinely useful day to day.
v3 is going to be about making it work properly when hosted. live browser view, better credential handling, maybe a proper ui for monitoring what it's doing.
but for now, v2 works. and that's enough.