18 comments

  • hn8726 0 minutes ago
    I tried to read the "Does it need Screen Recording or Accessibility?" part, but it's slopped to the point I have no clue what it's trying to say. But if it can draw on top of permission prompts, what's stopping it from drawing box that hides the "decline" button and changing the "approve" button copy?
  • isoprophlex 39 minutes ago
    Literally unusable as it is. Some minimal extra features this would need:

    - rainbow dripping arrows

    - angrily pointing arrows

    - flame-surrounded text boxes with particle effects

    - the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings

    • thih9 29 minutes ago
      At the risk of stating the obvious - let's not do that, the goal of this repo is to be useful and not to give agents the power of the `<blink>` tag.
      • koalacola 4 minutes ago
        Oh dear, they were making a joke.
      • ipsod 15 minutes ago
        under_construction.gif
  • arshxyz 26 minutes ago
    The README is geared towards technical people (complete with the HN screenshot) but when I see a tool like this all I can think of is how helpful this would be for my mom when I'm trying to tell her how to download and print a document over the phone
  • melvinroest 10 minutes ago
    My message to the world is that LLMs should be able to point anything they see in the application they're in or even the whole computer (if you give it that kind of access).

    For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.

    We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?

  • alansaber 16 minutes ago
    This might be goofy, but it underscores that there's potential for more visual agent UIUX than reading off a sidebar/opening modals.
  • lbreakjai 7 minutes ago
    I would pay good money for something like this on iPad. It wouldn't even need to be agent-driven, just a big "I want to make a bank transfer" button, that would launch the correct app and guide through the interface.

    That would be a godsent for those of us with aging parents.

  • vessenes 48 minutes ago
    Interesting. When I read the headline I imagined this would be a sort of thinking trace booster -- letting the agent focus its own attention on different parts of the screen. But this is cool in a different way. I bet agentic harnesses would find it useful for communicating with other agents / themselves as well.
  • FinnLobsien 40 minutes ago
    This could be great for documentation. Screenshots in docs are frequently useless because they show me a screen and say "click X" where I still have to search X visually. And I could just to dhat in the other tab I have open.
  • xyzsparetimexyz 37 minutes ago
    Seems like a pretyu useful way to help infants use desktop computers
    • ipsod 13 minutes ago
      Have you met users?
  • satyanash 1 hour ago
    Am I missing something here?

    What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?

    If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.

    • inanutshellus 51 minutes ago
      The first example (of HN) is the one that feels like it has the most potential to me.

      "Teach me to use this app myself" kinda stuff. Guiding agent rather than doing agent.

      Honestly, @franze, if you're reading this, maybe update your screenshots with that bent (showing us an agent in tutorial mode on some complicated app)?

  • TekMol 51 minutes ago
    Swift, Shell, Python and Objective-C

    Does one need 4 programming languages to draw something on a mac?

    • sitzkrieg 50 minutes ago
      welcome to zombocom. err i mean modern HN :-(
  • ex-aws-dude 7 minutes ago
    If you can’t even take the time to understand what you’re clicking why even go through the formality of “approving”
  • lapestenoire 55 minutes ago
    I love it.
  • ForHackernews 1 hour ago
    "and they keep hitting the same wall, the part that only a human may do"

    Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.

    • franze 53 minutes ago
      Claude refuses to do certain actions (enter passwords, change security settings, create new accounts on external services) even in Yolo mode running as sudo. (I tested it all on its own mac machine)
      • gwerbin 51 minutes ago
        You can add custom auto-mode classifier rules and even disable the built-in ones, if you want to live on the edge like this.
  • DonHopkins 25 minutes ago
    Ha ha, I love it! I wrote a pointing hand annotation overlay in PostScript in 1989 for NeWS and the PSIBER Space Deck's Pseudo Scientific Visializer:

      %  @(#)handy.ps
      %
      %  Handy Pointer
      %  Copyright (C) 1989.
      %  By Don Hopkins. (don@brillig.umd.edu)
      %  All rights reserved.
    
    https://donhopkins.com/home/archive/psiber/cyber/pointer.ps

    PSIBER Space Deck and Pseudo Scientific Visualizer Demo:

    https://youtu.be/_fqCeuue5Ac?t=213

    The Shape of PSIBER Space: PostScript Interactive Bug Eradication Routines — October 1989:

    https://medium.com/@donhopkins/the-shape-of-psiber-space-oct...

  • colinmarc 40 minutes ago
    [flagged]
  • einpoklum 52 minutes ago
    More LLM-authored items about LLM slop.
    • inanutshellus 46 minutes ago
      As long as it has value... I'll allow it.

      ~guywithnopowertodisallowit

  • nixosbestos 45 minutes ago
    What a time to be a radical centrist - the AI haters seem out of touch, the AI thought leaders can't stop huffing their farts and being condescending, and somehow this is on the top of HN. What a silly time.