Yeah, I'm not sure which way is supposed to be "sloppified". The default looks stylistically less slop-like, but obviously has the same filler content issue.
The blue one is the VibeTemplate_03. I see it everywhere.
The yellow one is at just a ripoff of an early 00s edgy news site. It could very well also be a VibeTemplate, but I've not seen a tool generate a site that looks like that by default.
For anything visual, they are still effectively blind, right? Still just working off the image embedding, sometimes using scripts to actually inspect individual pixels.
Blind seems too far here - Claude's increasingly able to diagnose visual issues on its own. I used to have to check every single image it generated, but now it's at the point where it'll catch most bad ones and regenerate them before it gets to the review step. Still misses some, though.
no. I just used Gemini to update a website I'm managing for a charity, and one of those steps involved Gemini, on its own, offering to scan the website for one of our beneficiaries, specifically the header images, to see if there was more information to include in our writeup about the beneficiary. And it did it quite well.
One way you could do it would be to find something that 1. would sound like a bug to a human and an LLM, 2. would be "confirmed" as a bug by a LLM, and 3. would be consistently solved in the same exploitable but reasonable way by an LLM. That's assuming a codebase that's largely AI-written with human review that you're able to open an issue for (possibly indirectly).
Another way is maybe something like saying there's a bug at some endpoint and thus manipulating a bunch of bots to DDOS that endpoint without having to pay for it?
Well no, Minecraft was already wildly popular while it was in beta. Luanti was explicitly inspired by Minecraft, as stated in the article it was named Minetest up until a few years ago.
>I guess no one actually wants to learn about harmony, about voicing, and voice leading, spend the hours. No one wants to learn how to actually play an instrument. I guess no one is willing to do the work. They want to just press some keys and declare that they made what the computer generated.
I don't understand that as a takeaway. Even if this worked perfectly, it would not be meaningfully stopping or discouraging anyone from learning music, and no one is claiming that what this models outputs is something the player played. It's a toy that someone made presumably because they like piano.
For every episode except the last 1-2 the characters are not textually aware that they're simulated humans made from a brain scan, and the conflict and plot are largely interpersonal conflict and being forced into adventures by the mad AI that runs the simulation. The characters wish that they could leave and/or change the simulation, and eventually accept their situation, and part of the conclusion is the idea that it's worth existing in their limited simulation and trying to make the best of it even though they'll never leave.
You could say it's exploring what it would be like to be trapped in a simulation with a mad AI, but I wouldn't say it explores the concept of ethics of simulating humans in general. Most of the things the humans don't like are arbitrary limitations of this simulation in particular (can't talk to the outside world, can't choose your form, can't have sex, can't have any impact on the world, forced into zany situations by an AI that doesn't really understand humans or listen to them), not anything about simulations in general. The finale touches on the idea of the simulated humans wondering about their original versions and what they're like and what they're up to, but that's about it.
Basically, I don't feel by watching the show I encountered any non-trivial ideas about human simulations just because it's about a simulation.
Yeah, the filtering process probably wasn't particularly robust. The Schrodinger's cat example ("It's a cat that has been misbehavin'!") sounds like... a joke?
Looking at the paper, it looks like they started with FineWeb-Edu, then filtered it based on an "age of word acquisition" dataset, with word frequency used as a proxy for values not in the dataset. They "only discard
samples in which more than 5% of the words exceed the target age of 12." Maybe 5% was too high? They also filtered out beyond K-5 math symbols, like sigma. Then they trained a classifier to do more filtering.
And they tested it on two grade-level benchmarks, and it only got 0-3% correct on the beyond k-5 boundary, while also decreasing in performance on the k-5 boundary (which they say is an acceptable tradeoff, since they were trying to get a sharp cutoff). So presumably since they got good results from the benchmark they stopped.
reply