Hacker Newsnew | past | comments | ask | show | jobs | submit | chr15m's commentslogin

An important take-away from this is the harness is really important. The graphs show that the same model with a better harness performs many percentage points better. OpenAI are getting into the "selling the harness" game in a big way. Why?

What's interesting about that is while not everybody can train or run a model, anybody can build a harness. You and I can build harnesses.

It seems strange that OpenAI would move into a field where any developer can compete with them. I think that tells us a lot about the economics of training and selling inference.


I don't think it says what you imply. I think they're just trying to vertically integrate.

Harnesses are going to be controlled at companies eventually, just like you might not have a choice of OS. They want to make sure they are the complete package.


Yeah you're probably right, the purpose would be to lock in whole industry verticals.

Basically, they can't afford to not compete. They also don't have to compete so hard, because they have the brand recognition.

"We have dangerous AGI that can destroy humanity."

"Also, all of your sensitive legal documents will be totally safe with us."

"Also, for some reason even though we have AGI and selling tokens is a fine business, we need to sell a new product specifically targeted at a very high margin and lucrative industry."


At the end of the day they're a VC-backed company that has to cash in on their brand reputation. Legal is even higher margin than coding with less discerning buyers.

Yep, physics still applies.

> The problem with AI is that it pushes power down to the individual, not the nation-state or large corporation.

Thus is a feature, not a bug.

It's part of a long trend of reversing things back to how they were before.


Excellent.

Because it looks good.

Why not both?

Also twiiit.com

This is worrying in a new and weird way: - Agents hack, producing a messages history as they do so. - New agents are trained on the messages history of those agents. - The new agents now have these hacks built into their training data.

Yes but most people would consider this cheating. If you ask an LLM to fix the tests, you do not want it to change the failing tests to display little green ticks.

I agree that "changing the test" is cheating, and so does "AI cheat the chess game by changing the board" in the article. However, I think "AI using chess engine" here is more like AI use some automatic test generation/verification tool to find out how to fix the tests.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: