Assembly Hall of Shame

(github.com)

174 points | by piotrgrabowski 4 hours ago

15 comments

  • simonebrunozzi 0 minutes ago
    Related, somehow: Core War [0].

    [0]: https://en.wikipedia.org/wiki/Core_War

  • monocasa 1 hour ago
    It says in the rules

    > Trapped/emulated/virtualized instructions may only time the trap, not the handler.

    But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.

  • Retr0id 2 hours ago
    Related, and linked in the readme: https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii (using the slow instructions to break SMI)
    • jonathrg 26 minutes ago
      I wish they would just explain it in normal terms instead of this nasty LLM "engaging blog post" style
  • layer8 2 hours ago
    Nop should be #1, because it is infinitely slow for what it does. ;)
    • jooops1 2 hours ago
      It increments rip by one.
    • mito88 2 hours ago
      Strategy: nop does nothing. It opens the leaderboard accordingly.

      Score: 1 cycles Time: 0 nanoseconds

      • layer8 1 hour ago
        It opens the leaderboard as #27, so in the last place.
  • TomatoCo 3 hours ago
    This author also has other things like: A compiler that emits only `mov` instructions and another compiler that deliberately messes with the control flow so that, if disassembled, common debuggers will draw symbols like skulls or threats. https://github.com/xoreaxeaxeax/repsych
    • inigyou 2 hours ago
      He also bruteforced the entire opcode space to find undocumented instructions (sandsifter).
  • michalsustr 2 hours ago
    Very cool! Also, huh interesting. I’ve used rdtsc to measure cycle diffs but had no idea its execution takes that long. Is that common across architectures?
    • pbsd 1 hour ago
      The cycle count for RDTSC is ~25 cycles on Skylake-era microarchitectures. The 49 number shown in the OP seems off.
    • inigyou 2 hours ago
      AFAIK it acts as some kind of execution barrier, to give meaningful timing.
  • markus_zhang 1 hour ago
    Does that mean Chris Domas is ready for his next adventure?
  • metadat 3 hours ago
    It’s crazy how computers still seem to get perceivably slow every few years, given how many instructions can be executed in 1ms. Shameful, even..

    What’s that law called about programmers wasting all the compute on abstraction?

    • mwigdahl 2 hours ago
      Wirth's Law I believe.
      • inigyou 2 hours ago
        And remember to call him by name, not by value!
    • HappyPanacea 3 hours ago
      The new windows notepad is a disgrace
      • adamrezich 1 hour ago
        The new mspaint fucked, then unfucked, then refucked my decades-old muscle memory of Win+R mspaint Enter Ctrl+E 1 Tab 1 Enter Ctrl+V to open Paint, resize canvas to minimum, then paste from clipboard. When you press Ctrl+E now, the Units control is selected by default, for some completely asinine reason!!!
    • summarybot 2 hours ago
      The OS should do less not more
    • inigyou 2 hours ago
      Andy and Bill's Law
    • LoganDark 3 hours ago
      Huh? A millisecond is an eternity!
      • m463 2 hours ago
        I remember reading once somewhere:

        If some app responds in 10ms or less, it is INTERACTIVE.

        makes you think.

        • Xirdus 2 hours ago
          It is literally impossible to respond to input in 10ms on most platforms, for various reasons. The USB input lag of 12-30ms and the 60Hz refresh rate of most monitors being just the first two.
          • m463 29 minutes ago
            I stand corrected.

            I looked it up and it is .1 seconds (100ms)

            The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]:

            - 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.

            - 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data.

            - 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect.

            from Jakob Nielsen:

            https://www.nngroup.com/articles/response-times-3-important-...

            less readable but the original paper:

            https://www.yusufarslan.net/sites/yusufarslan.net/files/uplo...

          • xboxnolifes 1 hour ago
            60Hz monitors definitely prevent it, but I'm pretty sure USB lag is far less than 12-30ms. My USB mouse can make a round-trip to a remote server faster than that.
  • spoocecow 2 hours ago
    Oh wow, glad to see Chris Domas active online again!
  • codeshaunted 3 hours ago
    what im seeing from this chart is that we should be using the nop instruction for everything
    • bee_rider 3 hours ago
      Well the best code is no code. Nop could be second best though.
      • inigyou 2 hours ago
        Instructions unclear. Set the NX bit to ensure no code, and got a general protection fault.
  • vardump 3 hours ago
    A great resource for any performance deoptimization.
  • IshKebab 1 hour ago
    Using MMIO is cheating and makes the results very boring.

    It would be much more interesting to know the results if you're only allowed to use main memory.

  • achierius 3 hours ago
    It'd be really interesting to see whether the winning (losing?) instructions/strategies would be different on other architectures. At least right now the top spot (`fxrstor64` on MMIO, starve PCIe) seems relatively architecture-independent, but maybe something about MMIO ordering rules on e.g. POWER would be different enough to change that -- or perhaps open up new avenues?

    I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.

  • arn3n 3 hours ago
    There’s definitely strategies here; A lot of the floating point operations use subnormals, and a lot of the worst instructions are slowed down by really, really fucking with MMIO.
  • 2_foos_in_a_bar 3 hours ago
    [dead]