Many people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it.
Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.
And while these sponsorship shenanigans are the tech business’s bread and butter, sponsored answer manipulation seems fundamentally more insidious. Even in a larger-scale measurement like this one, there’s no way to tell if any of that is sponsored, legitimately good recommendations, or the technical flavor of the goblins problem.
Any idiot can spend one minute now to acquire this amazing knowledge that you can just ask ChatGPT “I need a website”. If they don’t know now, sooner or later they’ll know from a short TikTok or something in the next couple of months/years. Not to mention AI companies are spending $$$ on ads to raise awareness; professionals already know they can create slides with LLMs, literally nothing’s different when it comes to most content websites. Yes, it’s an easy task any idiot can do. Head in sand, pretending it’s some sort of $10k for knowing where to hit situation isn’t going to reverse the trend.
I am pretty sure the average human would not have done this (among other slightly less absurd examples in the article requiring employees to fix it's mistakes):
Andon Labs added: “During the first week of operations, Mona purchased 120 eggs despite the café having no stove and to solve spoilage issues ordered nearly 50lbs of canned tomatoes intended for fresh sandwiches. Employees eventually created a shelf displaying Mona’s strangest purchases: 6,000 napkins, 3,000 nitrile gloves, industrial trash bags and 2.5 gallons of coconut milk.
> Was I going to reject the PR because of my code preferences? No, of course not. Ultimately it doesn't matter anymore.
You absolutely should've rejected it. Putting slop into production is a huge trap, because now the humans can't understand it and the LLM is just going to make it worse and worse over time.
They have a lot of moat, i'm not sure what youa re talking about. Only amatures are using Qwen, open source stuff that is 3-8 weeks behind. Plus OpenAI has some verticals that keep people in there.
I am not enthusiastic about criteria for human-like intelligence that imply that dyslexic people don't have human-like intelligence.
[EDITED to add:] I actually don't know whether dyslexic people find it difficult to count letters in words, if they have them already written down by someone else. I suspect they find it harder than people who aren't dyslexic. But perhaps "blind people whose spelling is poor" would have been better; I would not want to deny them human-like intelligence either.
> (assuming it actually follows the C standard and your compiler's documented extensions, which almost no real C code does - C was a horrible example) it will work just as well in a year as it does today.
> See how well it works to regenerate the same application next week, let alone next year. Prompts are fundamentally a different sort of thing from code
I mean... yes, but also no. C is actually a great example. So much of the code we wrote is about manipulating the specifics of that specific computer system we happen to be using at that exact moment. Everything from cpu specific instructions to how the ram behaves or how much of it there is all the way up to how library functions operate at any given point in time.
And in relatively short amounts of time, it can all change out from underneath you.
“Fun” fact, the rules were metaphorically written and blood, and grew to include things like “don’t microwave your clothes to dry them”.
Everyone who had worked in the rail industry (especially conductors and engineers) for more than a few years has horrible stories. PTSD is real. Suited by train is fucking selfish and disgusting.
We're not asking the model to simplify something, we're asking it to perform a task. Its subtle preferences show up as an overcomplicated path to the goal.
In some cases, there are also nuances that we don't pick up on. Here it's our preference for simplification that's showing up. We set the lossy compression factor higher than it does.
My comment wasn't very long, yet you somehow still ignored the main part, "and introduce new paradigms". The point is whether it can do everything humans can, entirely new theoretical frameworks and ideas, such as string theory or dark matter, are not coming out of AI at the moment.
Yeah let's start on a specific version that they overhauled the API since computers are getting new hardware features, god forbid a different DPI. And use that as a leverage point against a library that kept their API quite stable for the last 20 years!
Especially when comparing Qt against a library who couldn't keep its shit together for 5 years and is infamous for breaking all sorts of API and removing features.
i'm burning 500m tokens a day "writing code" for 83 days straight. one of those projects i'm building integrates netbox, stripe, quickbooks, mercury and deel all together via api. i wrote zero lines of code and read zero api docs. it does exactly what i want and gives me a 360 degree view of every aspect of my multi million dollar arr biz.
arc-agi3 is meaningless to most people. I'm not gonna look at the tests and see how hard it is. The actual test we look at is terminal bench, thats where software is being accelerated and closer to where rubber meets the road
I am so very, very tired of having to even consider, let alone make decisions about, the technical vs. political merits of products. It’d be so great if CEOs would keep their views to themselves.
Bro, "AI" can't even program decently. It's still, to this day, worse than every human I've ever worked with. The idea that it's going to take over every industry is laughable.
I got into an argument with a security guard at a gated community who tried to demand we give him a name, address, and "what's going on?" before he'd let us in.
Fixed that when I asked if he personally called 911 and when he said no, "So, before we make this a lot more formal, are you understanding that you are now obstructing emergency services personnel in the performance of duty?"
I by no means was suggesting that the trivial solution I, a non-mathematician, thought up in 20 seconds was somehow out of reach to a math professor who spent years on the problem. I knew I was wrong.
I didn't see why until I actually went through the solutions by hand.
Yeah, GPT4 was one-shotting utilities that GPT3 Davinci couldn't. So, I'd have my limited tokens on GPT4 crank out the initial program before iterating with my abundant, GPT3 tokens.
Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.