Your AI Agent Can Now Draw Arrows on Your Screen

Your agent stalls on a macOS permission prompt or an OAuth consent screen, and it has no business clicking either one. You walk over, squint at three overlapping windows, and try to work out which button it wants you to press. A small open-source CLI called bigarrow was built for exactly that moment, and it is useful in a narrower way than the hype around it suggests.
What bigarrow actually is
bigarrow is a small macOS CLI plus an agent skill that lets an AI agent paint big arrows, boxes, and text on top of your screen. The repo is franzenzenhofer/big-arrow-on-the-screen, and it's MIT-licensed. It ships as a skill for Claude Code and Codex.
Install is two commands:
brew install franzenzenhofer/tap/bigarrow
bigarrow install-skill
A few design choices tell you what kind of tool this is:
- No AI inside. There's no model, no daemon, no menu-bar icon, no account, no telemetry. It's a single small Swift binary. The intelligence is in your agent;
bigarrowis a pen. - Click-through overlay. It never steals focus, works on every display and Space, and drawing needs no macOS permission.
- It only points. A third-party summary at ai-tldr.dev states it never clicks, types, or captures the screen.
That last point is the one most coverage glosses over, so I'll come back to it.
The Hacker News thread (item 50018817) was posted by the author. I haven't checked the exact point total there, so I won't quote one. It got enough attention that people noticed, and that's all I'll claim.
The use case the README actually names
The README's core use case is handing a step back to the human. Its examples are macOS permission prompts, OAuth consent screens, 2FA, CAPTCHAs, passkeys, payment confirmations, signatures, and legal checkboxes. The agent points; the human decides.
That's a precise and honest framing, and it's different from the framing I've seen in the social-media pitch ("see what your agent is looking at"). Those are two separate problems:
| Problem | What you need | Does bigarrow help? |
|---|---|---|
| Agent hits a step it must not do alone (consent, payment, 2FA) | A clear visual handoff to the human | Yes, this is the design target |
| Agent acts on its own and you want to know what it saw | A record of the agent's perception | No. It draws, it doesn't capture |
| Agent runs overnight | Logs, saved screenshots, alerts | No. Nobody's watching the screen |
If you want "show me what you're about to touch" before an agent edits a record, bigarrow gives you the display half of that. The agent has to supply the target, and the agent has to be one that stops and asks. The tool doesn't make either of those things happen.
How an agent targets things
The targeting options are where the tool gets practical. An agent can point at:
- a screen coordinate
- a rectangle
- the mouse position
- a window
- an element by label (
--element, paired with--app) - a Peekaboo snapshot ID
Arrows expire on their own, which is a sensible default for something painted over your real work. point lasts 8 seconds and start lasts 300 seconds. An arrow also ends when the agent process that drew it exits, or when you run bigarrow stop. You won't end up with a stale arrow pointing at the wrong button an hour later.
The Peekaboo option matters if you already run UI automation. Peekaboo is a different kind of tool: a macOS CLI and menu-bar app for screen capture, accessibility inspection, and native UI automation, and it can be exposed to MCP clients. The pairing makes sense as a division of labor: Peekaboo perceives and acts, bigarrow communicates.
The README credits Peekaboo's visualizer and Peter Steinberger's Nameplate for the overlay window recipe, so this is built on known-good macOS techniques rather than anything exotic.
The accessibility catch with --element
Pointing at an element by its label depends on macOS Accessibility, and that comes with two gotchas worth knowing before you build a workflow on it.
The permission goes to the host app, not to bigarrow. You grant Accessibility to whatever app is running the shell: Terminal, iTerm2, VS Code, and so on. If your agent runs inside a different host than you expected, --element quietly won't find anything.
Chrome needs help. Inside web pages, --element only works if Chrome is launched with its renderer accessibility flag or VoiceOver is on. The README says this was verified on Chrome in October 2026. One caveat: if Chrome is already running, the flag may not be applied to it. Quit Chrome fully first, then relaunch it with the flag, and check the README for the exact flag and launch command.
If your small-business workflows live in the browser (invoicing portals, CRMs, banking), this is the setup step that will bite you first. Coordinates and rectangles work without it, but they're brittle: move the window and the arrow points at nothing.
Build requirements, if you'd rather compile from source: Xcode 16 or newer and macOS 14 or newer. The README reports 87 automated tests, with CI on macOS 15 and passing runs on macOS 26 and 27. That's the author's own reporting. Treat it as signal that it's maintained, not as independent validation.
Where it doesn't fit
I haven't stress-tested this in client work, and I haven't found anyone who has in a real invoicing or CRM workflow. So read the following as limits I can see from the README, not field results.
It's macOS-only. It runs only on macOS, building from source needs Swift and macOS 14+, and the skill is shipped for Claude Code and Codex. If your team is on mixed hardware or Windows, this isn't your tool today. I wouldn't claim it works with arbitrary agent frameworks.
It doesn't solve the "what did the agent look at" problem. Since it never captures the screen, an arrow tells you where the agent wants your attention, not what it perceived. If the agent misread the page and points confidently at the wrong cell, you get a confident arrow at the wrong cell. The tool makes the agent's claim visible. It does not audit the agent's perception. For that you want saved screenshots and diffs alongside it.
It does nothing unattended. An overlay only helps when a human is looking. For overnight automations, you need logged evidence and exception alerts.
The skeptics have a point. One Hacker News commenter asked why a big arrow is needed at all if the real problem is that confirmation dialogs don't get focus. That's a fair challenge: arguably the better fix is the operating system surfacing the right window. An arrow is a workaround for that, and a decent one, but it's worth being clear-eyed that it patches a UX gap rather than creating a new capability.
For contrast, there's the opposite direction: jimmyhmiller/overlay lets a human mark up the screen with boxes, arrows, freehand marks, and text, then exports an ordered PNG series for an agent to inspect. Agent-to-human pointing (bigarrow) and human-to-agent pointing (overlay) are two halves of the same conversation. If you're debugging an agent, the second one may be more useful than the first.
A five-second audit you can run without installing anything
The most useful thing in this whole topic doesn't require the repo. Take one workflow where an agent or automation touches something you care about: invoicing, lead records, client email. Then:
- Write down the single moment right before it commits a change.
- Ask: could I verify this in under five seconds just by looking?
- If all you get is a text summary, that's your weak spot.
The weak handoff is a message like "I'm about to update the customer record, approve?" The strong handoff shows the record, highlights the field, and puts the old value next to the new one. The second takes about three seconds to approve and catches real mistakes. The first gets rubber-stamped.
You can build the strong version several ways:
Weak: "Update customer #4471 email? [Approve]"
Strong: [screenshot with the field boxed]
Old: anna@old-domain.com
New: anna@new-domain.com
[Approve] [Reject]
A highlighted screenshot works. A before/after diff works. An on-screen overlay like bigarrow works too, if you're on a Mac and your agent is Claude Code or Codex. The requirement is the same either way: show me what you're about to touch, visually, before you touch it.
My own take, and it's an opinion rather than a sourced finding: the next gap in agents isn't smarter models, it's legibility. I'd take an agent that's a bit less accurate and visibly points at its target over a slightly more accurate black box, because I can catch the errors I can see.
Where this fits in the work I do
I build AI automations for small teams, and in many of them one step is left for a human to approve: the automation does most of the work, then needs a person for the one step that shouldn't be automated. My own approach is to design that handoff deliberately, so approval takes seconds instead of becoming a rubber stamp. Overlays like bigarrow are one possible presentation layer for that on a Mac. Saved evidence, diffs, and exception alerts cover the cases where nobody is watching the screen.
Should you try it?
Try it if you're on a Mac, run Claude Code or Codex, and regularly hit steps where the agent has to hand control back: permission prompts, consent screens, payment confirmations. It's MIT-licensed, has no telemetry or account, and the two-command install is cheap to undo.
Skip it, or at least don't rely on it, if you're on Windows or Linux, if your agent runs unattended, or if what you actually need is proof of what the agent perceived.
Before wiring it into anything real, check the license and setup requirements in the README yourself. I'd call it a promising building block, not a finished product, and I haven't run it against production workflows. The question it raises is the valuable part: for every step where an agent hands back to a human, can that human verify it in five seconds?
Want more like this?
Planning an AI automation project or need a second opinion on your architecture?
Connect with me on LinkedIn — Lazar Milićević, senior engineer who builds AI automations for solopreneurs and small teams.
Visit bizflowai.io for practical AI automation for solopreneurs and small teams.
Frequently asked questions
What is big-arrow-on-the-screen?
Big-arrow-on-the-screen is a project by developer Franz Enzenhofer that lets an AI agent draw overlays directly on your actual screen. The overlays include big arrows, boxes around elements, and text labels. They appear live on top of whatever you're viewing, rather than in a chat window or a screenshot sent back later. The idea is that an agent can show you what it's about to do instead of only describing it.
Why does AI agent oversight matter for small businesses?
Agents act on what they see, and when they get something wrong you often find out late: an invoice sent to the wrong client, the wrong row updated, or a report pulled from the wrong tab. The failure is usually that the agent's attention was invisible, so you couldn't check what it was looking at before it acted. Oversight lets a human catch mistakes before they commit.
How do I make AI agent approvals easier to verify?
Find one workflow where an agent or automation changes something you care about, such as invoicing or lead records. Identify the moment right before it commits the change, then ask whether you could verify it in under five seconds by looking. If you only get a text summary, require the agent to show its target first, using a highlighted screenshot, a before-and-after diff, or an on-screen overlay.
What is the difference between a weak and a strong human approval step in automation?
A weak approval step is a message like "I'm about to update the customer record, approve?" and tends to get rubber-stamped, letting mistakes through. A strong one shows the record, highlights the exact field, and displays the old value next to the new one. That version takes about three seconds to approve and is far more likely to catch errors before they are committed.
When should I use on-screen overlays vs logged evidence for AI agents?
Use on-screen overlays when a human is actively watching the screen, since they make an agent's intended target visible before it acts. For overnight or unattended automations, overlays do nothing because nobody is there to see them. Those workflows need logged evidence instead, such as saved screenshots and exception alerts. An overlay makes mistakes visible but does not make the agent correct.