Hacker Newsnew | past | comments | ask | show | jobs | submit | suprjami's commentslogin

The next board meeting, after buying at least 51% of the shares.

> or turn to alternatives like Firefox

Implying I ever left Firefox in the first place.


I’m a web developer and I have never dailied Chrome or a Chromium descendant. Firefox for daily, Safari for the occasional cross-check.

"If only there was someone could stop me, the CEO of Torment Nexus Inc, from creating the Torment Nexus."

> Coding is solved, perhaps, with unlimited token spend on a frontier model.

Coding is also a solved problem with unlimited salary budget on the very best developers.


Qwen 3.6 27B and 3.8 27B are the darlings of local inference at the moment.

The only other thing anyone is using is Qwen 3.8 Flash Next, only by memory-rich people.

Depending on which benchmarks you believe, these models (and the Ornith 1.5 finetune of Qwen 35B-A3B) are competitive at about Opus 4.5 to 4.7 level. That matches my experience in real tasks over the last few months.

Not bad for something you can run at home for a couple of thousand dollars.


the early 1.x versions of chad were tied to ornith. i still miss the speed of that moe. https://huggingface.co/nathansutton/Ornith-1.0-35B-UD-Q2_K_X...

I skimmed chad the other day. There's very little to it (by design). afaics you should be able to replicate what chad does by copying the chad system prompt into a SYSTEM.md for Pi. Then you could use it with whatever model/provider you want.

pi is a fantastic harness! they are a good default in the same way llama.cpp is. it works with everything and that is the point.

i was steering chad in the opposite direction. one model & one set of silicon -> taken to the max. swap out your CHAD_MODEL and it still runs, you just leave the drafter and the kernels behind.


Love the idea…

Except the "one model" is too small for a 64GB Mac much less 128GB, sad since the Q3 is proven less competent.

Offering a Q boost (with no leave behinds) on first run would be a bump worth some vibe coding while.


Updated the rootless podman container and restarted, worked fine.

Strange comment.

"If you remove the hype about how transformative they are, they really are transformative".

No. If you remove the hype about how transformative they are, you're left with what they actually are: a sometimes mildly useful tinker toy.


> The Astra pelicans are much better.

It's absolutely in the training data now.

Time to retire this (previously mildly amusing) benchmark as frontier models are now pelicanmaxxed.


What training data? Is there a big library of excellent SVG pelicans to train of?


Lots of SVG pelicans and lots of comments on what is good and bad about each.

Yes, all the people talking about this benchmark since OP created it.


Talking, sure. But have they created any good pelicans to train on? Pelicanmaxxing is clearly not a thing, which can be demonstrated by asking it for something else, or migrating to another domain (voxels, 3D modeling, and other things an LLM is supposed to be bad at, plus an unlikely combination of concepts). Looking at how relatively good for an LLM Astra is at Blender, it's obvious it was either purposefully trained for better spatial performance, or just generalizes better.

I actually think the pelican and similar tests are much less useful than it seems, but not because of pelicanmaxxing. They are supposed to serve as a vibe check of model's out-of-distribution performance, but do nothing to disentangle the generalization and memorization, which is the hard part. The combination of both is still useful though, and if you look at pelicans over time you'll see their quality is more or less correlated with that, with the exception of models specifically trained to produce vector graphics and 2D layouts.


This just lists every possible configuration permutation. Nothing insightful here.


So this is like OpenSCAD (code-based 3D modelling) but for music?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: