Hacker Newsnew | past | comments | ask | show | jobs | submit | timhh's commentslogin

Great read. Thanks for not using AI to write it! (Or at least making it not read like the usual slop.)


Thanks! Zero AI used to write it (:


I don't have a Cameo Pro 4 to test, but I have reverse engineered software to control some of their other devices. It's a little janky but if you want to help I'd appreciate it!

https://robocut.org/

I believe there is (or once was) also an Inkscape plugin somewhere but I never tried it.

As for how I reverse engineered it, I just used a USB sniffer and wireshark. I don't think there's a big risk of bricking these devices - I don't think they even have non-volatile storage, so worst case you turn it off and on again.

I would like to reverse engineer my solar power inverter too, but that is definitely brickable (the first one they shipped me came pre-bricked!) and is a lot more expensive so I've had to resist. :/


I'm still trying to convince them that RISC-V assembly exists...


I'm vaguely considering rewriting PAM in Rust. It's definitely something that would benefit from Rust's extra security and the code quality isn't that great. Nor is the UX. What "among other things" were you thinking of?


Implementation language is incidental. Rust is... fine, I guess? But I wouldn't mandate modules use it. What's important is to move PAM out of process and sandbox the process host (using landlock, namespaces like bwrap, whatever) so we can run sensitive code with least-privilege on both ends.


Yeah I always thought the whole architecture of PAM being a library is weird. A sane person would have made it a Daemon that processes talk to surely?

Really it should probably be part of SystemD, but I know that would anger the anti-SystemD zealots.


> A sane person would have made it a Daemon that processes talk to surely?

> Really it should probably be part of SystemD

https://github.com/systemd/systemd/pull/39855


It's confusing how they want a custom protocol that seems to be converging from JSON-RPC, just spelled differently. They should just use JSON-RPC. In any case, it'd essential to retain PAM's customization ability even if the bulk of the work is moved out-of-process.


Not everything has to be a "daemon". It's a bad trend in current programming zeitgeist to want to make everything into one. (sccache, no, you do NOT need a daemon to behave like ccache.)

For a PAM replacement, you want a separate process. It doesn't have to be a persistent separate process.


Is it valuable enough though. Looking at Google's stats Rust has several orders of magnitude fewer memory vulnerabilities even with `unsafe` (kind of the point). If C was at that level there's no way CHERI would have ever been proposed.

There are two counter-arguments:

1. There's a lot of C/C++ code still out there. You can't rewrite it all. I'm not totally convinced by that though because, a) do you need to? Google has shown that just writing new code in Rust is very effective, and b) AI is actually pretty decent at porting from C/C++ to Rust so maybe you can?

2. CHERI also allows really strong and fine grained compartmentalisation. This is absolutely fantastic for robustness, supply chain security and so on. If you want the absolute 100% most secure code possible, then Rust + CHERI with compartmentalisation is basically the best thing you can do. (Though Rust compartmentalisation is still not actually ready yet; it's in progress though.) That's really great but I'm not sure that level of security is needed by most projects, and also I think you can get pretty good compartmentalisation (though definitely not CHERI level) by doing something like what Xous does (basically isolation with processes/virtual memory, combined with the ability to call functions in other processes; IIRC Hubris OS does something similar).

CHERI is clever tech though and it would definitely be a boon for RISC-V if it succeeds.


The problem is that with LLVM, GCC, CUDA, Vulkan, POSIX, V8 and co, there will be lots of new C and C++ getting written as well.

Even on OpenJDK and CLR side, as new language features allow to rewrite even more runtime code from C++ into Java and C#, there is still new runtime code getting written in C++.

There is already clever tech for hardware memory tagging (SPARC ADI, and ARM MTE), CHERI is yet another way to tame unsafety on our computing stacks.


There's no real need for LLVM, GCC or CUDA to be memory safe. POSIX libc is of course C by definition but libc's are normally extremely well tested, and it is possible to avoid libc entirely if you want.

V8 is actually a nice case for CHERI since you can easily sandbox the JIT'd code. If you were to just rewrite V8 in Rust then you wouldn't get that benefit (you can't run the borrow checker on generated assembly). I assume they have some other sandboxing methods instead though.

But in general if you think about things like V8, that's used on high performance application class consumer CPUs. It's going to be at least 10 years before anyone has one of those with CHERI (unless ARM changes its mind about Morello). I would not bet against V8 being ported to Rust before that.

As I said I think CHERI is great technology and I hope it does succeed, but it does seem like the business case for it is not as strong as it was just a few years ago.


Why not? There are many cases where it matters, also routine CVEs show how well tested they are in practice.


Why not what? Sorry I'm not sure which bit of my comment you are responding to.


> Security through Obscurity still reigns.

That's not the case at all. The spec is developed in the open: https://riscv.github.io/riscv-cheri/

If you want to run CHERI code, it's true that silicon isn't easily available, but that's simply because it takes time. Various companies are working on it (Codasip, SCI, Secqai, lowRISC, etc.).

But you don't need silicon to run CHERI code. There are various emulators available that support it. There's QEMU: https://github.com/CHERI-Alliance/qemu There's also the RISC-V Sail model, this is the latest CHERI branch: https://github.com/CHERI-Alliance/sail-riscv (unfortunately it is a bit behind upstream master, and also a bit behind the latest CHERI spec which is still evolving).

There are also a few open source chips available that implement CHERI which you can run in Verilator or an FPGA. For example cheriot-ibex https://github.com/microsoft/cheriot-ibex . This is actually a variant of CHERI for microcontrollers called CHERIoT. Long story but the plan is to merge CHERIoT back into CHERI so it is just a "profile" of CHERI.


That looks quite simple. Think about something like this, a commercial SystemVerilog simulator (this only shows a fraction of the UI).

https://blog.reds.ch/wp-content/uploads/2018/09/questa13.png

Or something like Visual Studio.

Obviously most GUIs are not nearly that complex so immediate mode can get you quite far. Its biggest limitation is that it makes it hard to do some layouts. Your GUI layout becomes dictated by your data dependencies which is quite awkward.


Absolutely zero difficulty redoing this in a react style renderer. The only complexity is being careful with your data dependencies so as to not needlessly rerender.

Each pane is easily isolated, can share data with a view model scoped properly, etc. Writing it in an imperative toolkit is a "oops I forgot to update my data here" kind of hell. Data binding makes it slightly less worse


> Absolutely zero difficulty redoing this in a react style renderer

I would not classify a react style renderer as "immediate mode". It has aesthetic similarities with IM GUIs, but ultimately there is a fully retained tree under the hood, that gets diffed/mutated on every update


I like RISC-V (it's been my job for the last 7 years) but this is nonsense. Not everything RISC-V is good. CLIC was awful (thankfully it has been abandoned). The spec is not especially well written - the style is inconsistent due to being written by many authors, and it is waaaay too much of a textbook rather than a proper spec. (There is some ongoing work to improve this tbf.)

There's a practically unending list of undefined/implementation defined behaviours, which is great if you want to implement an ultra minimal microcontroller with 100 flops, but pretty awful otherwise.

Requiring the C (compressed) extension in the RVA profiles was definitely a mistake. The lack of true 16/64kB pages and conditional moves are probably a mistake (though fixable).

I don't know how any of these make it more robust and mature.

(But to be clear, I still think it's pretty good overall.)


I broadly agree with your points except one.

Requiring C (compressed) is necessary to avoid splitting the Linux ecosystem. Chips lacking C would never be able to run binaries compiled with C. There's no practical way for such binaries to detect this and work around it at runtime as they can with other extensions. And emulation would be super-slow given a large proportion of instructions are compressed.

Also the excuse given by Qualcomm - that it would make all instructions fixed length and so much easier to decode - is just wrong. RISC-V supports variable length instructions, even much longer than 32 bits, and you've just got to deal with it. Just because Qualcomm acquired a company with a microarchitecture that could only deal with fixed length instructions is no reason to break the ecosystem.

Also interested in the problems you see in Zicond. It claims at least to give you most of the benefit of conditional moves using only two instructions, but I've not actually tried using it. (https://docs.riscv.org/reference/isa/extensions/zicond/_atta...)


> Requiring C (compressed) is necessary to avoid splitting the Linux ecosystem. Chips lacking C would never be able to run binaries compiled with C.

Yes that's precisely the point of excluding it from the RVA profiles. It would mean that Linux distros don't compile code with C enabled, so chips are free to not support C and therefore can achieve higher performance (probably). And it opens 3/4 of the instruction encoding space for use by other things.

> Also the excuse given by Qualcomm - that it would make all instructions fixed length and so much easier to decode - is just wrong. RISC-V supports variable length instructions, even much longer than 32 bits, and you've just got to deal with it.

It's not wrong. RISC-V defines a mechanism by which 48/64 bit instructions might be used, but currently none are actually defined. All existing instructions are 16 or 32 bits. Without C all instructions are 32 bits. You don't have to deal with 48 bit instructions because there aren't any.

It's possible that they will add some in future, but I'm doubtful of that because a) it would be a huge pain, and b) they didn't for Vector which is where it would have been most useful.

> Just because Qualcomm acquired a company with a microarchitecture that could only deal with fixed length instructions is no reason to break the ecosystem.

Yeah it was too late to change but that doesn't mean it wasn't a mistake.

Zicond looks good - I forgot that exists.


> Yes that's precisely the point of excluding it from the RVA profiles. It would mean that Linux distros don't compile code with C enabled, so chips are free to not support C and therefore can achieve higher performance (probably). And it opens 3/4 of the instruction encoding space for use by other things.

The debate was between 16/32/48/64-bit instructions vs naturally aligned 32-bit and 64-bit instructions + new more complex instructions that require cracking to regain code size (things like load/store pair).

> RISC-V defines a mechanism by which 48/64 bit instructions might be used, but currently none are actually defined

The long-instruction-SIG just started a few weeks ago, and they are working on defining 48/64-bit encodings for instructions that could be used in future RVA profiles (so with high perf implementations in mind). If you are knowledgeable about this stuff, please get involved, so they don't mess it up. (not "you" specifically, but in general)


Compressed is necessary to reduce code size which is important for performance.

A bunch of vendors have done high performance server chips which support compressed (Rivos, Ventana, some Chinese vendors), so in actual reality this was only a problem for Qualcomm. And that's only because Qualcomm bought Nuvia and they wanted to do the cheap thing (minimally change the front end) rather than the right thing.


Of course you can make compressed work. E.g. you fetch 66 bytes instead of 64. Hell, Intel/AMD manage to make x86 fairly fast.

But it's definitely more awkward and has costs throughout the CPU.

I would be really surprised if the lower code density is worse than the improvement due to everything being nicely aligned. Especially because Qualcomm had actual data that it isn't (if you add new instructions with the extra coding space you free up).


You don't really fetch 66 bytes instead of 64, what real implementations do is read cache lines (of whatever size) and hold on to 2 bytes from the previous cache line if there was 1/2 a 32-bit instruction at the end of the previous cache line (the ISA has the 16/32-bit tag in the lower byte so you know how big an instruction will be even if you've only seen half of it)


> […] and hold on to 2 bytes from the previous cache line if there was 1/2 a 32-bit instruction at the end of the previous cache line […]

Well, and that is the worst case scenario from the performance standpoint since, if a 32-bit instruction is spans a page boundary, it will result in a page fault stalling the instruction decoder.

It might be acceptable in implementations not sensitive to such an overhead (e.g. embedded solutions) but is wholly unacceptable in high performance scenarios.


x86 seems to get by.

And how is an instruction spanning a page boundary and causing a page fault any worse than an instruction NOT spanning a page boundary and the next instruction causing the page fault instead?

As Paul said, if a 4 byte instruction spans a cache line/page boundary then you just hang on to the last 2 bytes of the page (first 2 bytes of that instruction) and decode them along with the instructions in that next cache line / page.

The only time it could possibly make a difference is if that spanning instruction is a jump to somewhere else AND that instruction could somehow have fit entirely in the previous page.

If there was no C extension then that next instruction would NOT be entirely in the previous page, it would be somewhere well into the next page, and that next page would have been required to be fetched much sooner. The C extension typically allows 30% to 50% more functionality to fit in each VM page.

Also, Qualcomm's proposed new instructions did not in fact use the freed-up space from not having C. They fit into other unused parts of the ISA.

I don't object to the new instructions Qualcomm suggested. I'd be perfectly happy to see them ratified and added to a future standard (even to RVA23 if they'd chosen to pursue that, but they didn't).

What I and others objected to was dropping the C extension from RVA23, or any future RVA-series, overnight given that RVA20 and RVA22 already existed with the C extension.

There will come a time when some RISC-V extensions will be retired and replaced, and it's entirely possible that C might be one of them, but there is currently no mechanism to do that, and when there is I'd expect that it would be done with a 10 or 12 year deprecation period, minimum.

NEVER overnight between one standard and the next one.

Which wouldn't have helped Qualcomm with their Nuvia core anyway.

Anyway, Qualcomm had now bought Ventana, which has engineers who know how to support the C extension with high performance, and they already had high performance RISC-V cores doing so.

So problem solved.


x86 is an ancient design mired in legacy and problems, so Intel/AMD Intel and AMD had to make it work for a modern world even if that involved kludges, band-aids, and crutches. Whilst you do have a point there, personally I do not find x86 particularly interesting to discuss.

I did not have Qualcomm and their latest bout of theatrical shenanigans concerning C in mind, either.

Back onto C.

  – C demonstrably improves code density;

  – Improved code density is conducive of better instruction cache utilisation and reducing the pressure on the TLB;

  – Mixed-width decoding (a trade-off) demonstrably adds front-end work (especially in pathological cases[0]);

  – No publicly benchmarked RISC-V processor is sufficiently contemporary and wide (or very wide) to impart the net effect of that trade-off at the highest performance tier;

  – Performance targets (aarch64 and x86-64) are moving faster than publicly demonstrated performance of existing RISC-V implementations in silicon, and the gap is not closing in.
So purported performance benefits of C may or may not materialise – it remains to be seen and proved, and claims that a future wide RISC-V core will validate design choices such as C remain not yet falsified projections rather than demonstrated engineering results – at this stage.

[0] The second page is not resident, the decoder can't proceed because the second half-word can't be obtained, the access generates a page fault, the CPU eventually vectors to the page-fault handler. Genuine page faults are extremely expensive.


The performance gap between RISC-V and x86 or Arm is in fact closing.

The SpacemiT K3 machines which most people who ordered in early May have now received are comparable to the RK3588 and Pi 5, with Rock 5 delivered to customers in mid 2022, Orange Pi 5 at the end of that year, and Pi 5 in October 2023.

So that's at most a 4 year gap, less than 3 years in the cast of the Pi 5.

That is the highest performance Arm64 machine most SBC users have. The faster CIX P1 exists but it seems that very few people actually have Orion O2 or Orange Pi 6 Plus.

Before the launch of the K3, the previous gen early 2022 to early 2024 JH7110, TH1520, K1 were somewhere around 6 years behind Arm SBCs.

RISC-V machines expected late this year (let's say early next year) will be similar to CIX P1, so just a 1-2 year gap.

Vs x86 the previous generation RISC-V was something like one of the last Pentium III or PowerPC G4, while the K3 is mid range Core 2 verging on early i5/i7 in many regards. So that's caught up on Intel by something close to ten years in four years. And the next gen will be somewhere around Zen 2 or whichever Skylake iteration is in the same ballpark.

Even more importantly than the gap, once that SkyLake to Zen 2 to Apple M1 performance band is reached, that is a performance level that remains "good enough" in 2026 for most users of computing devices for their everyday web browsing, media consumption, productivity/business app uses. I'm typing this on an M1 that I sit at all day every day (using it to access some faster machines for heavy work), and I have a lightweight Zen 2 laptop that I use for travel.


Performance gap is not closing in specimens commercially available today, and the promises of it «happening any moment from now» are now indistinguishable from monthly horoscopes.

For example, Zen 2 is a 2019 design, and the 64-core SG2044 C920v2, which was released in May 2025, a three-decode, four-dispatch core at about 2.6 GHz, with 128-bit vectors is still approximately[0]:

  – 9 times slower in the block tridiagonal solver at 64 cores;

  – 2.05 times slower in the lower-upper Gauss-Seidel solver;

  – 2.05 times slower in the scalar pentadiagonal solver.
Given a six year gap between two design (2019 vs 2025), the result is wholly underwhelming.

The reason why those three solvers are particularly interesting, especially in the HPC scenario, is because all three exercise substantial amounts of (unlike hobbyist and similar synthetic benchmarks):

  – Floating-point computation;

  – Memory hierarchy behaviour;

  – Cache utilisation;

  – Synchronisation;

  – Compiler optimisation;

  – Overall processor throughput.
Another, June 2026, study using production astrophysics codes found the SG2044 roughly 3–6 times slower[1] than an AMD EPYC 9554 system and 3–9 times slower than an Nvidia Grace system, workload depending.

[0] https://arxiv.org/html/2508.13840v1

[1] https://arxiv.org/abs/2508.13840v1


Huh?

There are RISC-V "performant" implementations on the best TSMC silicon process like x86-64 and aarch64?


There are not.


I thought so.

But I want that :)


> E.g. you fetch 66 bytes instead of 64

Not really, you would fetch fewer bytes with RVC [2, page 9], because the code density is better.

> I would be really surprised if the lower code density is worse than the improvement due to everything being nicely aligned.

This is very hard to quantify.

> Especially because Qualcomm had actual data that it isn't (if you add new instructions with the extra coding space you free up).

I've liked the back and forth slides bellow. Though I want to bring up to things regrading the Qualcomm slides:

    > [RVC] Performance benefit is modest
    > • Best case: 2-3% speedup
I recently benchmark compiling programs with a rva23 clang build and clang compiled for rva23-without-C and got a 10% performance improvement from RVC on the SpacemiT X100.

I also have no idea how they got those numbers. (not that they are wildly implausible, it's just not transparent)

    > Improving Android Code Size
In the last presentation they show how you can add a +-64M 32-bit long jump instruction to improve codesize in large binaries, like those in android.

I want to point out, that the JAL opcode has enough space left (7/8th) to encode a long jump of a same range and there is a proposal for a 32-bit +32M -12M 32-bit long jump: https://github.com/riscv/riscv-isa-manual/blob/zijfal/src/un...

[1] https://lists.riscv.org/g/tech-profiles/attachment/321/0/A%2...

[2] https://lists.riscv.org/g/tech-profiles/attachment/353/0/RIS...

[3] https://lists.riscv.org/g/tech-profiles/attachment/378/0/Res...

[4] https://lists.riscv.org/g/tech-profiles/attachment/400/0/AOS...


> RISC-V supports variable length instructions, even much longer than 32 bits, and you've just got to deal with it.

...no, not really? There is nothing like 9 byte-long MOVABS instruction of x64 that exists on RISC-V.

The main difficulty in decoding is that 32-bit instructions are not required to be 4-byte aligned, this means that naïve decoders will spend 2 cycles fetching such split instructions. It's possible to add a 4-byte ring buffer but all in all, efficiently supporting the C extension is non-trivial.


There are two options when designing an ISA to achieve competitive code size, add variable length instructions or add more complex fixed-length instructions which require cracking (2W instructions). The other option is: maybe codesize don't matter?

For high performance implementations both decoding variable length instructions and decoding/cracking fixed-length instructions into uops, are rather analogous in terms of the work hardware needs to do.

However, I think the advantage of fixed-length instructions, is that you can do further tricks, like pre-decoding in Icache. With RVC, you can also do pre-decoding, but now you need twice the amount of pre-decoding data, unless you find other tricks.

Still, in a reasonable variable-length ISA and fixed-length ISA, the variable-length one will get better code size. There are also a lot of other things to consider, RVC is self synchronizing, cracking is challenging for decode, but also keeps the backend better fed, how more instruction starts impact branch predictors, instructions crossing cache-lines...

I benchmark compiling programs with a rva23 clang build and clang compiled for rva23-without-C and got a 10% performance improvement from RVC on the SpacemiT X100. The X100 is a 4-wide out-of-order core and afaik doesn't do anything special for RVC, except for expanding the 16-bit to 32-bit instructions.

It's hard to quantify the real impact on a CPU design, but going the fixed-width route seems to enable more optimizations (not so much the decoding it self).


> maybe codesize don't matter?

Well, instruction cache still has limited size, and you still need to get your code into it. Paging in 512 KiB from the disk is faster than paging in 1 MiB from the disk.

> like pre-decoding in Icache.

I'm fairly certain x64 also does that?

> RVC is self synchronizing,

No, not really. You can still jump into the middle a 32-bit instruction, and it's possible it can be reinterpreted as a valid 32/16-bit instruction. Remember when people complained about how "overlapped instructions"/"hidden instruction streams" on x64 enable even more ROPs/gadgets than meets the eye? Don't worry, RISC-V has those too!


> No, not really. You can still jump into the middle a 32-bit instruction, and it's possible it can be reinterpreted as a valid 32/16-bit instruction. Remember when people complained about how "overlapped instructions"/"hidden instruction streams" on x64 enable even more ROPs/gadgets than meets the eye? Don't worry, RISC-V has those too!

Sure, it's a probabilistic thing. Whenever you see a 2-byte block starting with 0b11, you know there is an instruction start after those 2 bytes.

When I last looked at it, I got synchronize 99% of the time, when looking at a 8-byte block. In one qemu trace I empirically got that only 626 out of 70479 8-byte blocks don't contain synchronizations.

This certainly seems like something you could exploit quite well, if you are doing things like 8+ wide decoding. (maybe have a fast path that safes one cycle of latency)


RISC-V definitely does support instructions longer than 32 bits, starting at 48 bits (ie. 32 + 16), and going much longer. They are much easier to decode than x86 because the length is evident from the first byte. No ratified extension uses them now, but you're going to need to deal with them as the extension space gets more crowded. Including dealing with instructions split across cache lines and pages, and instructions aligned to 16 bits.

I'm not sure what point you're making TBH.


My point is that fixed-length instructions are supposed to be easier to decode than variable-length ones, right?

If not, then why even bother with fitting immediates and inventing LUI/AUIPC, just have a 48-bit long LI instruction. The same goes for 64-bit, an 80-bit LI.W is still shorter than the piecemeal construction with several instructions.

If yes, then the small cores are arbitrarily given a burden of supporting variable-length instructions, supposedly efficiently: if your instruction fetch is 16-bit wide, you need two fetches to fetch a single 32-bit instruction, which sucks; if it's 32-bit wide, you need to conditionally stash the upper half for the next fetch cycle, and still prefetch yet more 32-bits because that upper half may contain only a half of a full 32-bit instruction; alternatively, you can fetch 32-bits at alternated aligned/misaligned addresses and ignore the inefficiency of throwing away re-fetched bits — again, all of this sucks.


Yes, fixed-length instructions are easier to decode, but that also means a hard upper limit on the number of instructions that could ever be supported. Which is obviously a problem for a future-proof architecture.

The rationale for this and also for confining the base set to 32 bit is explained here: https://docs.riscv.org/reference/isa/v20250508/unpriv/extend...


I've read that rationale, and it's, well, I'm not going to say it's lying, but it's insincere. By the time it was written, they already settled on 16-bit alignment, and fetching (and then decoding) 16-bit aligned 32-bit instructions is either inefficient, or hard, or requires extra circuitry (or an instruction cache).


High performance RISC-V chips exist from Rivos, Ventana and others, and high performance variable length chips also exist in general (AMD, Intel). So in actual reality it increases complexity somewhat, but is not a problem.


sighs This kind of discussion is so annoying... I start with "RISC-V's decision to make 32-bit instructions only 16-bit aligned complicates instruction fetch and decoding", and the response is "For high performance RISC-V implementations, it's a trivial matter" (which is true), I response with "But then why even bother with mostly-fixed-but-not-quite instruction length, just make a properly variable length ISA, it'd even simplify the instructions", and the reply is "the low-end implementations would struggle with decoding that efficiently" — but they already struggle with instruction fetch when they implement C extension! Nah, it's fine, the high-end chips can cope with that.

And round and around this discussion goes... Apparently, RISC-V has no downsides at any end of the price spectrum, what a marvelous ISA.


If you have some specific question I may be able to answer it, since I've been involved with RISC-V for over a decade, dealt with all the main companies involved, written papers, etc. but so far I see no actual question or concrete objection.


[flagged]


You can see the encoding limitation it in the design on SVE, which only has destructive operations, but MOVPRFX, which is a round about way of doing 64-bit instructions, without doing 64-bit instructions.


When the top perf per watt or perf per MHz machine is RISCV, we'll talk. Until then, lol.


Many here wish for the non IP-locked RISC-V to get performant micro-achitectures for embedded/server/desktop/mobile that on the latest silicon process (without that, you can have a very good micro-architecture, that won't probably make the difference).

If some RISC-V high performance CPU manufacturers are being bought by big hardware actors: either they are scared of its competition and want to scrap it or they want to be part of RISC-V.


False dichotomy. They are also quite possibly aqui-hires which then put the good engineers unto actually useful projects.


This is another way to say they are scared ofe competition and want to scrap them.


Self-evidently, you therefore won't be the person building this machine.


Nope. I don’t want to work on systems that self-handicap in dumb and predictable ways. Mistakes happen. But mistakes that were preventable and knowable at design time … they are inexcusable, and RISCV is chock-full of those.


I disagree.


> small cores are arbitrarily given a burden of supporting variable-length instructions

Even the smallest commercial microcontroller cores e.g. the CH32V003, support the C extension. They strip out other things, such as half the integer registers, but they keep C.

And that's in a market where you can use literally any combination of extensions you want, because the customers compile all their own code, and you just tell them what ISA string to use.


English is not my native language and I wrote a bit too fast the message: I wanted to say that "everything pushing forward RISC-V is good".

I code RISC-V assembly, I don't use C machine instructions (I don't even use the pseudo-instructions, ABI register names and dodge nearly all ISA extensions, I try to stick to core as much as I can). I run my code on x86_64 linux with a small interpreter written in x86_64 assembly (thx to the 'R' in RISC).

I wonder if there are some 'broad and not niche, real-life' speed benchmark numbers to show how much C machine instructions are worth.

For the moment, I see those C machine instructions more as a marketing extension to match their arm equivalent: you know, for those key deciding people who care more about the amount of features and not their contextual pertinent usage.

To say an ISA is "good" is related to some set of technical sweet spots based on compromises based on projected usages.

RVA from my point of view is mostly preparing RISC-V hardware for some level of x86_64/arm compatibility.

I wonder if there are RISC-V implementations using the latest silicon process from TSMC.


Aarch64 dropped thumb instructions.

I don't think you're going to find a single benchmark on the effectiveness of compressed instructions since it really depends deeply on both the workload and the whole system. For example, memory bandwidth and cache pressure are both important for whether smaller text sizes matter, and that may depend on what else is running at the same time.

Note your assembler may be automatically compressing instructions without you asking. You'll have to disassemble the binary to find out.


You can benchmark stuff with and without RVC. Once there is faster hardware I want to do such a comparison with a full gentoo build for both sides.

However, quantifying what the result will actually mean is nearly impossible, because you don't know what the hardware cost was.


Maybe the best approach is to remove C from RVA while keeping it around in the specs for niche applications where text size _really_ matters (with current silicon processes, I wonder how weird those niche applications have to be to require C). But it seems some would remove C even from the specs to free some ISA space. If arm removed thumb...

Don't worry, I would know if the assembler is producing C machine instructions, my rv64 interpreter on x86_64 does not support the C instructions at all.


You can't remove C from future RVA because a large part of the value of RVA is that each version can run all the shrink-wrapped (binary distribution) code built for the previous versions.


Well, I said that because based on the documents provided here, it seems there are key people considering its removal even from the specs.

From my point of view, just do like all the others: clearly deprecate it, namely say that "from RVAx, don't create new machine code with the C extension". It is like in the linux kernel, it will then be removed very far in the future. But RISC-V is all about the far future, it can only be better to fix it asap.

After reading the comments and documents provided here, I am the first to be suprised by how much doing 'performant C' is not that easy and has a significant hardware cost. "arm removing thumb" should have been a strong signal.

Again, I am coding rv64 assembly almost every day, a good part could be C-ized to shrink text size, but based on various numbers provided here, why bother, better keep the 'R' of RISC as faithfull to its goal than anything else.


> "arm removing thumb" should have been a strong signal.

I don't see any reason to think that Arm ISA designers are any more skilled and knowledgable than RISC-V ISA designers. Given the relative number of people and the amount of academic and industry expertise you'd rather expect the opposite! Retired Arm ISA designers such as Dave Jaggar who designed Thumb and Thumb2 are on record as saying that the RISC-V designers have done a fine job, and that RISC0V is the current state of the art.

When I tested my Primes benchmark on a Pi 4, the variable-length Thumb 2 version was the fastest as well as, of course, the smallest.

    11.190 sec Pi4 Cortex A72 @ 1.5 GHz T32          232 bytes  16.8 billion clocks
    12.115 sec Pi4 Cortex A72 @ 1.5 GHz A64          300 bytes  18.2 billion clocks
    12.605 sec Pi4 Cortex A72 @ 1.5 GHz A32          300 bytes  18.9 billion clocks


But on the overall, based of the documents and information provided here, and what I think I managed to understand about them: it is not surprising some want to start the deprecation of C from RVA (some would even start its deprecation from the core specs to free some machine instruction encoding space).

It is a bit like compilers: 90% of the complexity/size is for 30% of the code optimization speed gain, which has mostly a significant and pertinent impact in specific work loads on our modern hardware. Basically it is "better" (since you get rid of 90% of the complexity/size of compilers) to combine pure assembly coding with less complex/bigger compilers ("10%") (and look at dav[12]d, not to mention ffmpeg, here the main issue is feature creep of ISO C new things and GCC extensions... and abuse of the nasm pre-preprocessor). For instance, with bzip2, with cproc/qbe (zero assembly though), I get 70% of the speed of modern gcc/clang. But what's really worrying: I talked (I think it was here on HN a long time ago) to one of TheHeavyThing assembly coders, and he said that with a brutal and naive hand-compilation of the corresponding C code, he got 15-20% faster than the best compiler could provide at the time on gzip compression. Something is up here, and I already did mention dav[1]d and ffmpeg.

We are in a 101 case of 'compromise', and here, it seems the more I read about it the more I think we are facing a case of 'benefits are not worth it in the end', well the apparent 'competent' people who are in the sauce seems to agree more and more about it, through the infos I have been getting from HN.

As for the 'compilation' benchmarks provided here, we know gcc/clang are heavily x86-64 oriented, namely goes into optimizations which should favor 'C' like machine instructions (destination register same than some source register), so the 1-2% "real" speed increase from some documents provided here sounds more like the real neutral/overall thing to me.


> CLIC was awful (thankfully it has been abandoned)

Could you expand a little on what made it a bad design? I'm not much up on RISC-V.


Mostly the specification was just poorly written with many ambiguities. But also the design was complex, weird, invasive and IIRC not backwards compatible with standard RISC-V.


> Default lock screen experience still has a needless delay of 5 seconds when entering a wrong (even blank wrong) password, even on the first attempt.

I suspect that is not KDE's fault (or Wayland's) - it's probably PAM, which by default has a 2 second delay (+/- 50%). That default is extremely difficult to change, but you can configure it. See my instructions here: https://github.com/linux-pam/linux-pam/issues/778#issuecomme...

Also if you follow that issue you can see I've been trying to convince the PAM developers to fix it (by changing it to a 0.5 second delay, which is much more tolerable and no less secure). Unfortunately they have this weird idea that users want the delay, because it lets them recompose their thoughts after getting the password wrong or something like that.


Wow, that has bugged me for years. Frustrating it's not easily configurable.


Good lord that thread is a dumpster fire. Thanks for finding out wtf is causing this, it has annoyed me for two darn decades, but never enough to go as deep into finding the cause…


Note the shunting yard algorithm is an iterative (as opposed to recursive) version of Pratt parsing (and also precedence climbing which is virtually identical). However as normally stated it does not do proper error checking - it will accept totally invalid input.

That isn't a fundamental limit of the algorithm though; you can easily add error checking, as I did here: https://github.com/Timmmm/expr/blob/1b0aef8f91460974d526b5ba...

I'm not really sure why this is omitted.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: