Ask HN: How come everyone is an LLM expert?

It is a bit strange how new model is released and hour after there is commentators declaring it complete trash and embarrassment to the AI industry, or the best thing since sliced bread. Surely they have not had the opportunity to test the ins and outs of the model yet? Or do people just blindly trust benchmarks as if they were not pretty easy to manipulate, as research has shown quite a few times now? Or is it just all vibe?

So how do you measure how one model is better than another?

2 points | by delis-thumbs-7e 1 hour ago

3 comments

  • spottedmarley 44 minutes ago
    I built my own benchmarking arena that tests local models on all of things the I need a model to do well. I don't look at any of the existing benchmark data that is out there. When a new model drops, I run it through my arena and see how it compares to previous models. If I talk about a model being good I am referencing my own accumulated knowledge on how a model performs for me on tasks that I care about. I generally will never be heard talking negatively about a model (except maybe a frontier/hosted model, they all suck in their own ways) because if a model sucks it just gets deleted and I move on to other things. I suppose I'd consider myself somewhat of an 'expert' when it comes to analyzing local model performance, but I don't really listen too much to what anyone else says about them, or which benchmarks tell them which things about a model. Just test them on the things that are important to you.
  • tolugenius 1 hour ago
    There is probably far, far more people trusting benchmarks and "I remade x thing in 1 prompt with y model" claims than you'd imagine, just ignore all of it. You know your workflow and what better should be and could be, measure on what works for you. You should note (and I may be wrong, not active in these part) a lot of those demos are very toy, recreating a known game, known app, known workflow, etc. very interesting but again a very toy example that should be taken with that in mind.
  • bigyabai 1 hour ago
    Benchmarks can still be useful, even for benchmaxxing labs. For example, the recent Beam model benchmarked much worse than DS4.1 Flash and GLM 5.3, both of which perform extremely well outside of benchmarking. Regardless of whether or not Beam was benchmaxxed, it's performance suggested that it wasn't capable of solving problems that other models in it's weight class could do easily.

    The best-case scenario is that Reflection was being honest about their model's disappointing performance. The worst-case scenario is that they benchmaxxed, and it still managed to underperform compared to it's peers.