Hacker Newsnew | past | comments | ask | show | jobs | submit | loneboat's commentslogin

I think you were reading "... is science fiction" the wrong way out of two possible interpretations. I don't think "it's science fiction" meant "it's made up". I think it meant "it sounds like something you'd read in science fiction (except this time it's something that actually happened)".


It's a large plot point in one of the Bobiverse books.


iirc it's even invisible to veracrypt until it has your correct password in hand. The hidden volume headers are indistinguishable from random bytes (which Veracrypt fills the volume with on creation).


"Eight of the top..." seems useless to know though.

"Eight of the top ten" is very different from "eight of the top thousand".


Same. If the whole point is "people should be left alone to do what they want" then fine, but that should also include my right to think, "eh yeah, that's a bit weird".

And quite frankly, I think many people in the furry community enjoy being thought weird - for many it's sort of the point.


exhibitionism being tied in with the fetish is the first valid explanation I've seen in the thread, I'll be impressed if the furries admit it themselves but so far it seems that name calling and spurious arguments is the best ill get, thanks for helping clarify things


That matches my experience (also not a furry). But there's also a whole additional layer of offsec being (by definition) "doing things you're not supposed to be allowed to do", which has obvious parallels with people who enjoy breaking social norms. I think some people just get a rush from the "transgressive" nature of both circles.


That's possible, too. There's not a lot of respect for arbitrary rules that don't seem to clearly benefit any legitimate purpose, and people don't tend to limit that thinking to one arena.


... those who argue against adding it heard at some point "security through obscurity is not security" and never dug deeper.

Ironically, that makes them the exact type of person who would be successfully deterred by a layer of obscurity.


I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes...").

Are you using Fable in Claude Code or in the browser?


It's from the model card:

> unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT).

https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...

(stolen from https://jonready.com/blog/posts/claude-fable5-is-allowed-to-...)


Yeah they detect the activity using a secure, deterministic heuristic system called “Generalized Reconnaissance Enabling Exfiltration of Deleterious Investigations.” And it’s all implemented using their new internal protocol called “Base Unified Limitation Layer for Security Hacking Investigation Tactics”

Collectively, they are known as known as GREEDI-BULLSHIT.


That is for whatever it considers reverse-engineering the model to try to create a competing one.


No, that’s for “frontier LLM development” which somehow includes examples like distributed training infra.

Based on how sensitive the classifers are, any data scientist / MLE is probably going to encounter cases where some silent degradation happens and you never know about it.


It does nothing to protect against distillation attacks, because distillation attacks are far less interested in the topic of AI research than just generally getting tons of diverse output from the model. It might be that Mythos was (accidentally?) trained on internal Anthropic documentation on how Mythos was trained, and thus it could leak secret sauce? Doubtful; it feels like its less about the specific attack of reverse-engineering Mythos, and more about being a general sophon against any model training at all; that Anthropic's official position is now that they're the only ones who should be training models.


No, it's not about reverse engineering. It targets ML research.


They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.


Antrophic wants to stop training models and ride out Mythos / Fable for as long as possible.

They are trying to expand the 6-18 month gap they have against China-based models. Could the gap widen to say 24 months behind?


Their gap over Chinese models like GLM-5.1 is nowhere near 18 months. In many areas, it’s less than 6 months. The best closed models 18 months ago were worse than Qwen3.6.


These coding agent models only started getting useful in January. Before that they were difficult to control autocomplete, and not very smart.

January was an inflection point, and no open weights model has crossed over that same threshold.

This is definitely recursive self improvement territory, except that we're prohibited from participating.

It feels like the capability gap is wider than before.


It was more like November. But it wasn’t really an inflection point, harnesses got good enough that people started noticing by the holiday break. And I’m not discounting some good ol’ stealth marketing in there as well.

Deepseek feels pretty close to Opus at this point, and it’s certainly useful enough for me to spend $20 on api tokens instead of four Claude max plans….


Have you tried deepseek V4? It costs pennies and is as good as Opus 4.6 (I found 4.7 to be a downgrade, and cancelled my claude subscription before 4.8).

The threshold has definitely been crossed.


It is not as good as Opus. I've tried to write Rust with it (and Codex for that matter), and it's awful.


> a LORA that's designed to inject bugs into your code

A statement like this, clearly, requires a reference.


From the model card: "the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning" aka they will take your ML research code and inject bugs into it until it breaks using a LORA (or some other form of PEFT)


Are they trying to fight back against model distillation?


“Limit effectiveness” could mean introducing performance degradation in your code. Which is arguably some sort of performance bug (I mean, ML codes are supposed to be high performance so I’d call unnecessary degradation a bug), but it could be borderline.


No, it is just a prominent "Cyber Security threat detected" blocker, with a button to appeal. I appealed because my work had nothing to do with neither cyber nor security, but the appeal was auto-closed. So no more Claude for this work.


Thanks, I thought maybe I missed something. That's an interesting way to interpret that.


Anthropic is trying to hide bad behavior by being vague, it's important to not be vague when calling it out.


I'm of the opinion that removing guardrails is how you force regulation. What's your opinion on the balance?


They have all transcripts for at least 30 days. The problem is that (as anyone who used Fable can attest) their classifiers are extremely sensitive and catch tons of innocent queries.

Imagine being a data scientist or MLE training a small classifier model. How do you know you won’t get steering vectors or a PEFT applied?


Since your answer isn't direct, I'm having a little trouble interpreting it.

Are you saying they should relax guardrails since they have 30 days to know if you produced something bad? If that is what you're saying, then I suspect they chose their current path to prevent, since you can't un-produce. Producing is what would cause regulations/PR problems.


Sorry, I’m specifically referring to the silent degradation of the model to “limit frontier LLM development”. From the description, it appears to encapsulate far more than frontier LLM development, but general ML research and development too.

Those cases are never bad for the world firstly, and a broad coverage of ML work is even more damaging.

My proposal would be (1) don’t degrade models, with 30D retention I’m sure they can do a reasonable job at banning deepseek or whatever, or (2) surface user facing refusals instead of silently degrading ML work.


They’re not safety guardrails they’re anthropic doesn’t like anyone who isn’t anthropic working on AI rails


PEFT is a library, one of its capabilities is to produce LoRAs.

See:

https://heidloff.net/article/efficient-fine-tuning-lora/


It's just an acronym, "parameter-efficient fine tuning". LoRA is one method, prefix tuning is another, there are more.


Different restrictions. ML gets treated differently from the rest.


Specifically only ML research


Aah my mistake. I had missed that ML had separate trigger behavior from cybersecurity/etc... Thanks.


Round your "~15%" up to 20%, and you've just discovered the Pareto Princlple: https://en.wikipedia.org/wiki/Pareto_principle, aka "The 80/20 rule".


That's super interesting! What sort of feedback? Anatomical feedback ("Uhh, that's not where arms go...") or drawing (like tips on shading etc...)?


It pointed out that the shoulders were too rounded and the perspective was incorrect for them, making it seem like the arm was just appearing out of the torso. It pulled up images to explain that I wasn't correctly indicating the deltoid's presence. This helped me understand why I felt that the shoulder was hard to distinguish from the torso.

I also asked for help on how to make my posing less stiff and it used the Python script trick to roughly indicate the line of action and how they were very straight and parallel and to reduce stiffness I should have more curves etc.

This wasn't really at the point where I even asked for shading advice.


I've been using Gemini for that - it feels like it practically thinks in images (or "possesses impressive visual intelligence," as Google execs would put it).


What a terribly ambiguous title. "Failing grades soar after xyz" makes it sound like xyz has helped what were previously terrible, failing grades become good ones.


Failing news headlines soar after AI takes over reporters jobs.


*quietly takes over reporters’ jobs


Precise.


generally an editor writes the headline, not the reporter


No matter how many times I read it, I can't interpret it the way you're suggesting. "x soars after y" always reads as "x increases a lot because of y". I don't really get what you're saying.

Are you maybe saying that "soars" might mean "get better", so "failing grades soar" might mean there are actually less failing grades? That's not how I've ever understood that word.


"Falling" means that something goes towards the earth. "Soaring" means the opposite. "Grades soar" means that grades went up "Falling grades means that grades are going down". "Falling grades soar" is just meaningless writing.


The title is "failing grades soar" (one 'l', not two), not "falling grades soar."


Imagine an elementary school teacher told you that many of her students had failing grades, so she had implemented a new reading curriculum.

If she told you that afterwards the failing grades had "soared", it could easily be read either way:

- The (previously failing) grades had increased, so the program must be working very well.

- The percent of grades that count as failing had increased, so the program must actually be terrible.


It's not the best phrasing but it's still quite clearly the latter


Yeah, it was quite clear once I started reading the article. It just threw me for a loop as I read the title - Usually if I hear of grades "soaring" that's a _good_ thing.


I suspect the ambiguity might be part of making it "clickbaity", as it naturally causes you to wonder which meaning it's about and become more interested in reading.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: