Hacker Newsnew | past | comments | ask | show | jobs | submit | nichochar's commentslogin

my read is that PG is one of the lucky people who was very successful before kids, so he could just put YC in self-driving mode by handing off to the next generation and his life's work still feels meaningful.

I think it's very different if you have not achieved your life's work yet.


Healthy things. For me its riding a bike, which is magical because i) its solo activity so my brain meditates ii) its healthy but no shocks on my knees and body iii) its technology free, no notifications on a bike


these people have ai too. AI increases competition amongst the ambitious, so it doesn't solve the problem


i really like your framing, this encapsulates the problem very consisely:

> Time is finite. Desire / drive are not. Solve for X


Thank you! That's very kind. It's definitely not an easy equation to solve and is a constant moving target :)


(original author)

this is true for me, wouldn't be remotely possible without my wife.


yep all of this resonates


Surprisingly pragmatic and info packed article..

Kudos to databricks, I also find it interesting that such different companies (Stripe, Ramp, Databricks) are all building the exact same internal tools.

I think building companies is going to look more generic in the future because intelligence is an API now.


Thank you for the feedback. We wrote this because after discussing with some of our peer companies, I realized everyone was roughly doing similar things. And I thought it would be good for someone to just systematically write down what those are so that others can try out the techniques if they find them useful.


+1 well written, well paced article. Pleasure to read.

Have you tried measuring Gemini? now that you have the router it should be a simple task. Thanks!


yep, that's what i qualify as "not an advantage".

It's totally fine if thats what your company is like, but the labs have an unfair advantage.


This was often true when writing code manually to be fair.

You could get to "something that works" rather fast but it took a long time to 1) evaluate other options (maybe before, maybe after), 2) refine it, 3) test it and build confidence around it.

I think your point stands but no one really knows where. The next year or so is going to be everyone trying to figure that out (this is also why we hear a lot of "we need to reinvent github")


When I hire fresh out of college… I can see them coming in and not having the slightest comprehension of the difference of the things that they did in school to get a grade and never touch it again versus a product that is supposed to exist and work for 10+ years.


We ran some tests at mocha (we have a coding agent with our own harness to build web apps, with a lot of tools and medium length tasks (3min to 10min).

Our notes:

Sonnet 4.6 feels like a fundamentally different model than Sonnet 4.5, it is much closer to the Opus series in terms of agentic behavior and autonomy.

Autonomy - In our zero-shot app building experiments, Sonnet 4.6 ran up to 3-4x longer than Sonnet 4.5 without intervention, producing functional apps on par in terms of quality to the Opus series. Note that subjectively we found Opus 4.5 and 4.6 are better "designers" than Sonnet 4.6; producing more visually appealing apps from the same prompts.

Planning / Task Decomposition - We found Sonnet 4.6 is very good at decomposing tasks and staying on track during long-running trajectories. It's quite good at ensuring all of the requirements of an input prompt are accounted for, whereas we were often forced to goad sonnet 4.5 into decomposing tasks, Sonnet 4.6 does this naturally.

Exploration - In some of our complex "exploration" tasks (e.g. cloning/remixing an existing website), Sonnet 4.6 often performs on par or better than Opus 4.5 and 4.6. It generally takes longer, and takes more tokens, though we believe this is likely a consequence of our tool-calling setup.

Tool-use - Sonnet 4.6 seems eager to use tools; however, we did find that it struggles with our XML-based custom tool use format (perhaps exclusive to the format we use). We did not have a chance to assess with native tool use

Self-verification - Similar to Opus 4.5/4.6, Sonnet 4.6 has a proclivity for verifying it's work.

Prompting - We found Sonnet 4.6 is very sensitive to prompting around thinking, planning, and task decomposition. Our prompt built for sonnet 4.5 has a tendency to push sonnet 4.6 into incredibly long thinking and planning loops. Though we also found it requires significantly less careful and specific instructions for how to approach problems.

How are we thinking about this:

We can't launch this model day 0, it requires more changes to our harness, and we're working on them right now.

But it reminds me a bit of 3.5 to 3.7 --> It's a pretty different model that behaves and responds to instructions in new ways. So it requires more tuning before we can extract its full potential.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: