v1 was amazon q getting eyes. v2 was those eyes getting sharper, 109 tools, network interception, stealth mode, the whole thing. v3 is different.
v3 is about you and the agent sharing the wheel.
the problem with watching
when i hosted the mvp on a server, the thing that broke the experience wasn't the tools. it was that i couldn't see what was happening. the browser was running somewhere on a vps, doing things, and i was just reading logs.
so the first thing i built for v3 was a live dashboard. a web ui that streams the actual chromium browser in real time using mjpeg, basically a continuous jpeg stream at around 60fps. you open the dashboard in your browser and you can watch amazon q navigate, click, fill forms, all of it live.
that solved the visibility problem. but then i thought, if i can see it, why can't i touch it.
the interaction overlay
on top of the live stream there's a transparent overlay div. it captures all your mouse and keyboard events and maps them back to the actual browser coordinates.
the coordinate mapping was the tricky part. the stream is scaled to fit your screen, so a click at position 400, 300 on your display might actually be 640, 480 in the real 1280x800 chromium window. the overlay does that math on every event so what you click is what actually gets clicked in the browser.
what you can do from the dashboard right now:
- click anywhere on the page, left, right, middle button
- double click
- scroll with your mouse wheel, mapped to the exact position
- type into any focused field, printable characters are batched at 60ms for efficiency
- press special keys, enter, backspace, escape, arrows, tab, all the function keys
- keyboard shortcuts, ctrl+c, ctrl+v, ctrl+a and so on
- mousemove, throttled to 60fps so hover states and tooltips work
so you can literally fill out a form. click the input, type your text, tab to the next field, fill that, hit enter. all from the dashboard while amazon q is watching the same browser.
hand in hand, not hand off
the thing i want to be clear about is this isn't you taking over from the agent. it's both of you in the same browser at the same time.
amazon q can be navigating to a page while you're watching. it hits a captcha, it pauses and surfaces a modal asking you to handle it. you solve the captcha directly in the browser through the overlay, submit it, and the agent picks up from there. no copy pasting, no describing what you see, you just do it.
or the agent fills a form but gets a field wrong. you can just click that field and retype it. the agent sees the corrected state and continues.
that back and forth is what i mean by hand in hand. the agent does the repetitive navigation and data work, you handle the things that need human judgment or credentials or just a quick correction.
the command bar
there's also a command bar at the bottom of the dashboard for when you want to drive without touching the overlay. you can type commands like:
- a url to navigate directly
- scroll down, scroll up, scroll left, scroll right with optional pixel amounts
- click followed by a css selector
- type followed by a selector and text to fill a field
- js followed by any javascript to run in the page context
- back, forward, reload as quick chips
the command bar is useful when you know exactly what you want to do and don't want to aim at the stream. the overlay is better when you're exploring or reacting to what you see.
the activity log
below the stream there's a live activity log fed by server-sent events. every tool call amazon q makes shows up there in real time, the tool name, whether it's running or done, a summary of what it did. you can see the agent working even when it's doing things that don't produce visible changes in the browser, like reading the accessibility tree or checking network requests.
the log also shows the status dot in the top bar. green when idle, amber and pulsing when the agent is actively running a tool.
where this leaves things
v3 is still being built out. the dashboard works, the overlay works, the stream works. what i want to add next is a way to annotate the stream, like draw on it to point the agent at something, and better session management so you can save a browser state mid-task and come back to it.
but even as it is now, the experience is completely different from v1. v1 was me explaining to a blind agent what was on the screen. v3 is me and the agent looking at the same screen, both able to touch it, working through things together.
that's the version i actually wanted to build from the start.