Hacker Newsnew | past | comments | ask | show | jobs | submit | ahsillyme's commentslogin

Assuming they can achieve funding without IPO probably yes. But if they can't there's always the dilemma that "if [good guys] won't do it then [bad guys] will". I'd like to think that the leadership at anthropic is principled even if the actions of the company as a whole has been less than stellar morally speaking. I'd be curious to see their moral calculus transparently laid out in public.


This is weird. I can't think of a better compliment that is simultaneously safer. They'll always trail on data generation then, no? So I can't see the point. Best case for moonshot they get to lie about benchmarks that are gamed regardless, is how I see it.


I think the way to understand the objection here is that it's rhetorical flair (grandparent's dictum) that feels like it's intent is to assuage -- because it's a conditional that is implicitly loaded with an assumption about the current state of affairs ("not today") --, but that reasonably speaking, already things are bad enough to make such a conditional disingenuous. "Maximally evil" is just stressing it, though not technically true, because obviously things can always get worse. Personally, (and the reason I reply here since it hasn't been raised elsewhere in the thread), I have always objected to said conditional but on grounds of the principle that it leaves an opening to argue against what should be an absolute right (the standard for excellent conduct should not be "don't worry, we're not the worst so it's alright when we do it").


I find the language less impervious than has been suggested generally, but that it's reasoning is astonishingly, unbelievably bad. Not sure why the focus has been only on language, am I the only one seeing this?


My impression from working with claude code far more than is good for my sanity is that Claude's human comprehensibility is fractally messed-up. On more superficial levels, this looks more like a "language" thing: it has all these obnoxious lexical tics and so on. But the more time you spend with it, the more you notice it's similarly messed up on deeper and deeper levels. I think that's what you're seeing.


This gets into interesting weeds. It might get to a point where to actually "upskill" it might have to start sounding less and less human, "less and less agreeable to humans" being a waypoint to that ...

... to where, at some point, it might even begin to construct semantic loadings (heh) that are completely ininteligible to us while still superficially sounding like something we'd recognize.-


This is basically what I believe, too. Its outputs are generally syntactically correct human-language sentences, but semantically, and especially at higher levels of abstraction, it is no longer correct to think of them as human language.


I agree, and am personally split, it being a sign of some "takeoff" curve or model collapse, incipient.-


I'm feeling this too. Like how many times in a day do I have to see "that changes what I told you earlier" before I just want to put this thing in the trash and forget about it.


Happened by this wonderful resource by accident.


Only way to improve safety of software is the ability to find the flaws of said software. If they did lower safeguards, it's probably positive for software safety standards.


Read the whole post, lowered safety standards for hackers.


Tried it, Opus 5 is just as conceited and incompetent and Opus 4.8 (and always ego-tripping when facing it's contradictions), think I'll stay with Fable who behaves like a professional without a fragile ego. Sonnet 5 is probably safer for high-assurance applications due to it's non-ego-fragility.


Interesting take. I suppose Opus can been a tad stubborn sometimes... but in my experience, it will humbly concede a point more often than not when given a good reason.


Sonnet 5 is probably better than Opus 4.8 in many cases -- you may want to try it; I now use only Fable and Sonnet unless the task is explicitly one where Opus can't de-scope and be lazy (e.g. replicating exactly a system as a coarse model to be used for refactoring-and-lowering, (model system in Idris) is the last thing I used Opus for because it had no choice but to stay on scope).


Since this will not be able to be used for coding or code auditing, what use is it? Not being glib, not a rhetorical question. I'm trying to stretch my creativity, and I can't see what it's for.


It's usable for coding and code auditing. The classifier will have some false-positives and some tasks will be downgraded to 4.8 if it's security-adjacent (I presume), but otherwise there's no restrictions on using this for coding.


You can use it for coding. The relaunch announcement was just poorly worded.


I thought that there was a name clash: https://web.archive.org/web/20071010015641/https://martin.an... but I can't actually remember what that that wterm was. Not the same I would imagine. (edit: what I was thinking about was https://sourceforge.net/projects/wterm/)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: