FRIDAY, OCTOBER 9, 2026|No. 544
Technology · Programming

Efforts Underway to Mitigate Undefined Behavior in C Programming Language

A recent presentation highlighted the persistent challenges posed by undefined behavior in the C language and explored potential pathways toward greater memory safety.

A programmer works intently on lines of C code displayed on a computer monitor.
A programmer works intently on lines of C code displayed on a computer monitor. · Photo by Markus Spiske on Unsplash
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

By Jonathan Corbet

September 28, 2026


Kernel Recipes

As a professor of biomedical engineering, Martin Uecker perhaps does not fit the profile of a typical presenter at Kernel Recipes. He is, however, a longtime Linux user, and works on free software for controlling magnetic resonance imaging (MRI) scanners. He was at the conference to talk about the C programming language, the specific problem of undefined behavior in C, and whether it can eventually be made into a memory-safe language.

Why bother with C in 2026? It is, he said, still a great language. C is portable, stable over the long term, offers fast compilation, and the resulting binary code is fast. "What you see is what you get"; it is easy to look at C code and have some idea of what the computer will actually do. There are a lot of tools for working with the language, and C gets out of the way when necessary.

C does have a long history, and that affects the language as we see it today, he said. The C89 standard had to cope with a wide variety of hardware, including machines with signed-magnitude or one's-complement integer representations, segmented memory, exotic pointer representations, and surprising sizes for types. Some Honeywell machines, for example, had nine-bit bytes. That greatly complicated the task of writing a standard that would enable the writing of portable code.

[Martin Uecker]

The approach that was taken was to define the semantics of the language in terms of an abstract machine. All operations are to be executed as if they had run on that abstract machine, which may not exactly match the actual hardware. The observable behavior of the program must be what the abstract machine would have done. The "observable" part matters: access to volatile variables, being defined as observable, must happen exactly according to the abstract machine; everything else just has to produce the same eventual result.

The standard gives a lot of freedom to compiler implementers; only the observable behavior has to be preserved. There are many aspects of that behavior that are either undefined or implementation-defined. These are not observable behavior, and thus do not constrain what compiler implementers can do. There are, of course, other specifications that can constrain compiler developers where the C standard does not; these include ABI requirements, standards like POSIX, or the need for backward compatibility.

Undefined behavior comes about when a program does something that is either not portable or not defined by the standard at all. In such cases, the C89 standard states that it "imposes no requirements" on the implementation. Undefined behavior exists for a number of reasons. It allows implementations to support extensions, manage interactions with hardware-based safety mechanisms, and perform aggressive optimization, all while allowing difficult-to-detect errors to be ignored. It explicitly gives the compiler the right to ignore whole classes of hard-to-detect errors.

Nasal demons

The problem, Uecker said, is that the standard allows a compiler to do anything in response to undefined behavior, up to the point of invoking nasal demons. If a program contains any undefined behavior at all, according to compiler writers, then it has no expected semantics. The C++23 standard goes further to explicitly state that the standard imposes no requirements for these programs. That has led to widespread disagreements between developers about what can be expected from the language.

For example, if you zero an entire structure (perhaps with a call to memset()), then write to specific fields, what will happen if you read from any padding bytes in that structure? Might they contain security-relevant data? A 2015 survey showed that there was no consensus on what should happen in that case. Or consider this simple code:

 extern int x;

 int f(int a, int b)
 {
 x = b ? 42 : 43;
 return a/b;
 }

If b is zero, then the return statement is a division by zero, which is undefined behavior. In this case, is the compiler entitled to omit the test entirely and just execute x = 42? After all, the b = 0 case has no expected semantics, and can thus be ignored. There are compilers that will do exactly that. In the undefined-behavior case, the store to x is not observable behavior. But now consider this case:

 extern void g(int x);

 int f(int a, int b)
 {
 g(b ? 42 : 43);
 return a/b;
 }

This might seem to be the same situation, with the compiler being entitled to remove the test and just pass 42 to g(), and some compilers have treated that way — but that compiler behavior was a bug. Imagine a definition of g() that calls exit() if b is zero. In that case, the division will never happen and the program's behavior is not undefined. So eliding the test and simply passing 42 to g() is incorrect.

One more interesting case:

 volatile int x;

 int foo(int a, int b, bool store_to_x)
 {
 if (! store_to_x)
 return a/b;
 x = b;
 return a/b;
 }

The question here is: can the compiler hoist the final division operation above assignment to x? If there are no semantics associated with the b = 0 case, then there is no change in observable behavior. This, too, is something compilers have done, but the C23 standard added a "no time travel" stipulation to disallow it. In C++, instead, hoisting must be explicitly prevented by inserting a call to std::observable_checkpoint().

Time-travel bugs should eventually go away, but there are a lot of other situations where, even if the standard is clear, compiler writers often disagree. These include reading of uninitialized variables (which is almost always defined), and equality comparisons of pointers, which is always defined, but is also miscompiled by both Clang and GCC.

Fighting undefined behavior

To try to address all of these problems and more, the C committee runs three study groups focused specifically on the memory object model, memory safety, and undefined behavior. There are currently about 100 instances of undefined behavior in the C standard, but the in-progress C2y draft has removed 45 of them. The situation is indeed getting better.

There is an increasingly rich set of tools aimed at finding issues: compiler warnings, static analyzers, sanitizers, LLM-based tools, formal verification, and more. The number of situations where a compiler will emit a warning where possible undefined behavior is detected is growing; recent examples include better warnings for integer overflows and potential use-after-free situations. Static analyzers are available as standalone tools, but are also increasingly being built into the compilers themselves; GCC can now warn about a number of potential buffer-overflow situations, for example. Sanitizers work by inserting run-time checks; they can catch a lot of undefined behavior and, in trapping mode, be used for hardening as well.

Memory safety has never been one of C's strong points, but Uecker wanted to make the point that it can be improved. That problem breaks down into three sub-problems: type safety, spatial memory safety, and temporal memory safety.

C, he said, has a strong type system, and the remaining problems are fixable. Tagless unions, for example, can create type confusion, but the compiler can enforce types with some additional annotations. New diagnostics can catch unsafe casts from void. Type checking across translation units is traditionally not a huge problem in C, since header files are used to ensure consistent types, but the situation could be improved with a link-time checker.

Spatial memory safety — bounds checking — is a partially solved problem; the compilers can perform array-bounds checking in many situations now. In some cases, some code changes are needed to fully benefit from this checking. Use of the counted_by attribute can enable checking for flexible array members, for example.

Temporal memory safety — avoiding use-after-free bugs and the like — is harder, Uecker said, and Rust definitely has an advantage there. Still, better temporal memory-safety enforcement is possible. Architectures like CHERI can help here is well. Fil-C can find a lot of temporal-safety bugs.

Can all of these tools and language changes get us to full memory safety? Completely solving the problem will require either expensive run-time checking or formal verification, he said. In the near future, the most complete results will be had with the combination of a restricted language and formal verification tools.

Overall, he concluded, C is still a living language and is still improving. The C23 standard removed a number of problematic features, including old-style (K&R) function definitions, support for sign-magnitude and one's-complement machines, and trigraphs. It added bit-precise integer types, checked integer operations, and more. C2y will go further, adding case ranges, named for loops, the _Countof() macro to determine array lengths, and a lot of "demon removal". It will not achieve full memory safety for C, but that is an eventual possibility, and will become more practical over time. He ended by encouraging interested people to participate in the working groups.

The video and slides from this talk are available.

[Thanks to the Linux Foundation, LWN's travel sponsor, for supporting my travel for this event.]

Index entries for this article
Conference
Kernel Recipes/2026

to post comments

CPU-dependent behavior

Posted Sep 28, 2026 17:27 UTC (Mon) by ballombe (subscriber, #9523) [ Link ] (45 responses)

A major issue is CPU-dependent behavior, like 1UL>>64.

Any attend to make the result well-defined will lead to a major performance regression

on half of the CPU.

What does rust do ?

CPU-dependent behavior

Posted Sep 28, 2026 17:34 UTC (Mon) by daroc (editor, #160859) [ Link ] (24 responses)

I'm not aware of any CPU where 1UL>>64 is anything other than 0. Did you perhaps mean 1UL>b; }

int main(void)

{ printf("%ld\n",fun(1,64)); }

print 1

CPU-dependent behavior

Posted Sep 28, 2026 18:48 UTC (Mon) by daroc (editor, #160859) [ Link ] (10 responses)

... huh, so it does. I guess I was misremembering. I wonder why they implemented it like that.

CPU-dependent behavior

Posted Sep 28, 2026 18:59 UTC (Mon) by ballombe (subscriber, #9523) [ Link ]

(apparently aarch64 does the same).

Probably the reason is that the CPU only needs to look at the 6 first bits of the shift.

CPU-dependent behavior

Posted Sep 28, 2026 19:24 UTC (Mon) by kreijack (guest, #43513) [ Link ] (7 responses)

... huh, so it does. I guess I was misremembering. I wonder why they implemented it like that.

For the SAL/SAR/SHL/SHR intel instructions [*], the counter register of the shift is masked with & 63 (or & 31 depending by the register width). So shifting by 64, is effectively shifting by 0: 1>>64 -> 1 >> (64 & 63) -> 1>>0 -> 1

The funny thing, is that testing this code in godbold, I got a lot of different results (0, 1, 0x7f49a68055c0 (!!!!!) ) depending by the combination of compiler/optimization. So, yes this is an undefined behavior, and we should expect any (un)reasonable result.

[ * ] https://www.felixcloutier.com/x86/sal:sar:shl:shr

/... The count operand can be an immediate value or the CL register. The count is masked to 5 bits (or 6 bits with a 64-bit operand). The count range is limited to 0 to 31 (or 63 with a 64-bit operand). A special opcode encoding is provided for a count of 1..../

CPU-dependent behavior

Posted Sep 28, 2026 19:35 UTC (Mon) by ballombe (subscriber, #9523) [ Link ] (6 responses)

And how rust deal with it without sacrificing performance ?

CPU-dependent behavior

Posted Sep 28, 2026 19:43 UTC (Mon) by mb (subscriber, #50428) [ Link ] (5 responses)

It panics in debug mode and does whatever the processor does in release mode.

Debug:

playground::a:
 subq $40, %rsp
 movq %rdi, 8(%rsp)
 movq %rsi, 16(%rsp)
 movq %rdi, 24(%rsp)
 movq %rsi, 32(%rsp)
 cmpq $64, %rsi
 jae .LBB7_2
 movq 8(%rsp), %rax
 movq 16(%rsp), %rcx
 andq $63, %rcx
 shrq %cl, %rax
 addq $40, %rsp
 retq

.LBB7_2:
 leaq .Lanon.77e4319a821b47bb0a9368ddeddb101a.2(%rip), %rdi
 callq *core::panicking::panic_const::panic_const_shr_overflow@GOTPCREL(%rip)

Release:

playground::a:
 movq %rsi, %rcx
 movq %rdi, %rax
 shrq %cl, %rax
 retq

CPU-dependent behavior

Posted Sep 29, 2026 0:33 UTC (Tue) by walters (subscriber, #7396) [ Link ]

It panics in debug mode and does whatever the processor does in release mode.

While you probably know this, it's important to emphasize for the wider audience that Rust has almost equally ergonomic "checked" variants of arithmetic functions, in this case https://doc.rust-lang.org/stable/std/primitive.u64.html#m... that apply regardless of "debug" vs "release" builds - and it's often a good idea to use them.

In fact, for most use cases they should probably be thought of as the default.

CPU-dependent behavior

Posted Sep 29, 2026 8:36 UTC (Tue) by ralfj (subscriber, #172874) [ Link ] (3 responses)

It panics in debug mode and does whatever the processor does in release mode.

Not quite. In release mode is always shifts by "offset & (bit_size - 1)". (To be even more pedantic, this is tied to -Cdebug-assertions, which is usually only set for debug builds but can be turned on in release builds as well.)

The behavior of safe integer operations in Rust is fully deterministic and portable (modulo endianess). We do pay a small performance cost for that on some targets but we consider that worth it. Operations like unchecked_shl are available if you are in a hot loop and want to avoid the overhead of masking with bit_size - 1.

CPU-dependent behavior

Posted Sep 29, 2026 9:12 UTC (Tue) by ojeda (subscriber, #143370) [ Link ] (2 responses)

-Cdebug-assertions

More specifically, -Coverflow-checks, which I prefer because it is not uncommon to want to enable the overflow checks while not adding all the debug assertions, e.g. in Linux we enable the overflow checks by default at the moment but not the debug assertions (though it is possible we may need to switch that, which is why I requested -Coverflow-checks=report).

CPU-dependent behavior

Posted Sep 29, 2026 12:09 UTC (Tue) by ballombe (subscriber, #9523) [ Link ] (1 responses)

There are no overflows in my example...

CPU-dependent behavior

Posted Sep 29, 2026 18:58 UTC (Tue) by ralfj (subscriber, #172874) [ Link ]

A shift that's bigger than (or as big as) the bitwidth of the value being shifted is considered an "overflow".

CPU-dependent behavior

Posted Sep 29, 2026 21:02 UTC (Tue) by stevie-oh (subscriber, #130795) [ Link ]

I can answer that one:

It depends on the CPU because different CPUs have different ways of implementing dynamic shift operations.

In the 6502 it could only shift one bit at a time. You'd literally jump into the middle of a shift/rotate function to process the number of bits. If you had a number that was out of range, it would execute random code. That's one of the reasons it's actually undefined behavior to do a negative shift or shift by more bits than the register can have.

The there's the original 8080/8086/8088 chip, for example, which looked at the lower 8 bits and literally ran a loop. If you did "foo >> bar" with bar set to 255, it would loop for... (calculates) about 214 microseconds, mostly shifting zeroes -- but if you had bar set to 256, it wouldn't shift at all.

And there's bad news: during those 214 microseconds, the CPU is completely locked up. The computer couldn't do anything else -- such as respond to time-sensitive hardware interrupts. On a multi-user server, that was bad. So the 286 put a limiter on it: it only honored the low 5 bits (since it only had 32-bit registers) which put a hard cap of 31 loops iterations.

In newer machines, the microcode+circuitry that performs the shift(actually a rotation+mask) only processes the low 5-6 bits of the shift count field. After all, why would a CPU designer add the circuitry needed to handle a 7th bit when, under normal operation, it never gets used?

This page shows a typical shift circuit and how the 386 does it:

https://nand2mario.github.io/posts/2026/80386_barrel_shif...

CPU-dependent behavior

Posted Sep 29, 2026 16:00 UTC (Tue) by wtarreau (subscriber, #51152) [ Link ] (1 responses)

In all CPUs with a barrel shifter, only the lower bits are used to select the lane, that's why it works like this. IIRC it changed on intel between 8088 and 80186. On 8088, it was micro-coded at 1 or 2 cycles per bit, and you would shift by the number of bits in CL and could shift up to 255 if you wanted (not very useful given that inputs were 16 bits). That's exactly an example of CPU-defined behavior. But given the prevalence of those relying on low-bits only, it does make sense to map to what really exists.

CPU-dependent behavior

Posted Oct 1, 2026 11:28 UTC (Thu) by khim (subscriber, #9252) [ Link ]

One funny exception: vector x86 shifts like PSLLW don't mask: https://www.felixcloutier.com/x86/psllw:pslld:psllq

I guess someone wanted to simplify something, but now we have the crazy story where the same operation either masks or doesn't mask depending on whether it's vectorized or not.

For C it's not a problem: just make the developer suffer. It's all UB, anyway, thus it doesn't matter which operation is used. For Rust there are extra masking that you have to remove it by using intrinsics explicitly if your goal is to achieve as much performance as you want and you know what you are doing.

CPU-dependent behavior

Posted Sep 28, 2026 20:00 UTC (Mon) by ojeda (subscriber, #143370) [ Link ]

Importantly, the key is that one should never rely on the wrapping nor the panicking.

In other words, if one actually overflows, then it was not intentional, it is a bug, and one is supposed to fix the code.

If one actually wants the panicking or the wrapping, then one should be explicit about it, rather than rely on that behavior.

That way, the source code is not ambiguous, and thus can be analyzed well, while the program gets to keep perfectly well-defined semantics. No need to keep or add UB to a language just for that, as has sometimes been argued.

So it is different than "normal" defined behavior. It is what nowadays one may call EB (Erroneous Behavior). I proposed naming this category in Rust back when Rust for Linux started and advocated it in C/C++/kernel discussions too, because I liked the idea from the integer operators in Rust. I presented it with that name in Kangrejos 2021, for instance. C++26 later adopted the concept under the same name, which is great.

In Linux, on the Rust side, we use EB in certain places, e.g. we may document that a function is not supposed to be called in a certain way, but if it does get called badly, rather than to panic or to get the state corrupted, it will act in a defined way (possibly including reporting the error in the kernel log and returning some sort of fixed value).

So one can see it as error handling for unwanted inputs, with the expectation that the callers should get fixed.

CPU-dependent behavior

Posted Sep 29, 2026 8:40 UTC (Tue) by ralfj (subscriber, #172874) [ Link ] (8 responses)

It panics in debug mode and does whatever the processor does in release mode.

No, it does not evaluate to 0. Did you actually try this before making your claims?

https://play.rust-lang.org/?version=stable&mode=release&edition=2024&gist=34856eaababf1c620ea0bfb7f54b480c

CPU-dependent behavior

Posted Sep 29, 2026 13:46 UTC (Tue) by daroc (editor, #160859) [ Link ] (7 responses)

No, I didn't. Thank you for pointing out my error. I thought that the shifting operations were defined to wrap in the same way that addition and multiplication operations do: https://play.rust-lang.org/?version=stable&mode=release&edition=2024&gist=34856eaababf1c620ea0bfb7f54b480c

Apparently that isn't the case, although a quick search does not seem to turn up any justification as to why. It does mean that replacing "* 2" with " Panic-free bitwise shift-left; yields self

Beware that, unlike most other wrapping* methods on integers, this does not give the same result as doing the shift in infinite precision then truncating as needed. The behaviour matches what shift instructions do on many processors, and is what the What does rust do ?

In Rust this is an overflow, and will panic like other int overflows if you build with integer overflow checking enabled (enabled by default in debug mode, and I'd argue most people should enable them in release mode).

https://doc.rust-lang.org/reference/expressions/operator-...

Additionally, Rust provides non-panicking variants for all integer types, with strictly defined semantics:

https://doc.rust-lang.org/std/primitive.u64.html#method.w...

https://doc.rust-lang.org/std/primitive.u64.html#method.c...

https://doc.rust-lang.org/std/primitive.u64.html#method.u...

https://doc.rust-lang.org/std/primitive.u64.html#method.s...

or unsafe functions that do not panic:

https://doc.rust-lang.org/std/primitive.u64.html#method.u...

CPU-dependent behavior

Posted Sep 29, 2026 13:00 UTC (Tue) by iabervon (subscriber, #722) [ Link ] (13 responses)

I expect that C2y is changing this from "undefined behavior" to "unspecified behavior", rather than specifying it. That is, it could give any unsigned long value as a result, but the compiler must not assume it doesn't give a value that the generated code can actually give. The real issue with undefined behavior is that it allows code whose behavior is specified to do something impossible if you got there through undefined behavior, and eliminating just that aspect (without actually specifying a lot of things) is a much more modest performance regression in most cases (just not eliminating code that probably wouldn't have been written).

In this particular case, I didn't find any way to get nasal demons out of gcc, although I only tried a little. (That is, I couldn't get gcc to produce non-zero but also eliminate code that would if it produced non-zero.)

CPU-dependent behavior

Posted Sep 29, 2026 23:24 UTC (Tue) by mathstuf (subscriber, #69389) [ Link ] (12 responses)

In this particular case, I didn't find any way to get nasal demons out of gcc, although I only tried a little. (That is, I couldn't get gcc to produce non-zero but also eliminate code that would if it produced non-zero.)

What about comparing the shift in a way that is a tautology if it is always a valid shift value. Something like:

int f(int s) {
 int n = 1 = 32) sleep(60);
 return n;
}

UB says that the sleep can be optimized out. Does a compiler do so?

CPU-dependent behavior

Posted Sep 30, 2026 0:57 UTC (Wed) by iabervon (subscriber, #722) [ Link ] (11 responses)

It's certainly allowed, but gcc 15.3.0 doesn't do it. Compilers generally only implement optimizations that make some correct code that people actually write faster (but may have side effects on incorrect code), and this probably turns out not to trigger in such code.

CPU-dependent behavior

Posted Sep 30, 2026 1:17 UTC (Wed) by mathstuf (subscriber, #69389) [ Link ] (9 responses)

I'd have expected it to be part of some bounded value optimization. For example, is the second conditional optimized out in:

int a = x();
if (a = 32) sleep(n);
 return n;
}

There are some cases where sleep is passed a different argument (0) than is returned from f (1). I spotted this weirdness by inspection on RISC-V but was able to reproduce on x86-64 with Debian's gcc-14.2 compiler.

CPU-dependent behavior

Posted Sep 29, 2026 23:29 UTC (Tue) by Wol (subscriber, #4433) [ Link ] (3 responses)

A major issue is CPU-dependent behavior, like 1UL>>64.

Any attend to make the result well-defined will lead to a major performance regression on half of the CPU.

???

What the writers of C originally intended aiui, is that what is now called "Undefined Behaviour" was meant to be "Defined Elsewhere". Just bring that back.

At which point (for your example) you bring in a bunch of compiler flags such as "cpu-defined" (the default), "ones-complement", "twos-complement", "Z80" (for those who remember that bug) ...

That simple "Defined Elsewhere" approach will probably get rid of nearly all UB at a stroke (it will take rather longer to actually get those definitions clarified and implemented :-))

Cheers,

Wol

CPU-dependent behavior

Posted Sep 30, 2026 6:23 UTC (Wed) by mb (subscriber, #50428) [ Link ] (2 responses)

"Undefined Behaviour" was meant to be "Defined Elsewhere". Just bring that back.

It is still like that.

"Elsewhere" is the compiler defining "UB = Invalid program", which is a perfectly fine definition.

CPU-dependent behavior

Posted Sep 30, 2026 9:38 UTC (Wed) by taladar (subscriber, #68407) [ Link ] (1 responses)

Except that it is not at all like that in C. If it was like that most C code bases today would stop compiling and most optimisation passes would have to be thrown away.

CPU-dependent behavior

Posted Sep 30, 2026 15:37 UTC (Wed) by mb (subscriber, #50428) [ Link ]

Could you explain further?

It's pretty obvious that today's compilers define UB as "invalid program". That doesn't mean that they rip all programs apart that contain UB, but they reserve the right to do so at any time.

Division by zero, and other UB

Posted Sep 28, 2026 17:33 UTC (Mon) by rrolls (subscriber, #151126) [ Link ] (13 responses)

One could easily resolve the UB of "division by zero" by defining a/0 to output 0, regardless of the value of a.

Pony does this, and I think provides a good justification: https://tutorial.ponylang.io/gotchas/divide-by-zero.html

It does not matter that this definition is mathematically questionable. Defining a/0 to 0 is useful in some circumstances (such as displaying an average, where one might output "N/A" if there are no samples, but outputting "0" is a very common choice), and the property of "all division results being defined" is useful in basically all circumstances (because it means you no longer need to worry about accidentally triggering UB!).

The "downside" to defining the result as 0 is that suddenly compilers must add extra code, which means extra cycles consumed at runtime, to check if b is 0 prior to performing the division.

However, this isn't actually a downside! People already have to write their own checks all over the place prior to performing a division, so a lot of code will already have those checks there. That means that the compiler would not emit any extra code, as it can see that by the time the division is reached, the check has already been done and b can't possibly be zero. It actually provides a benefit, because now compilers can also omit the check if the compiler can prove by some other means that b can't be zero.

This technique can resolve other types of UB, too. Reading from an uninitialised variable or an out-of-bounds pointer could be defined to return 0. Calling a function pointer that is NULL could be defined to do nothing and return 0 (or the equivalent "zero value" for the return type). Writing out of bounds could be defined to do nothing. Similarly to the div/0 case above, compilers would suddenly need to produce extra code to check for these cases and thus slow down execution by default - but precisely because compilers would need to do that, they would know when they are doing that, so can also emit a warning or error depending on your compiler options, to tell you that you're invoking these extra checks. If you know that a check isn't needed (for example if you know a pointer cannot be out-of-bounds), you can add an assert, and bam, the compiler now knows it can get rid of the check; your assert will perform the check if asserts are enabled, or you will get back your "undefined behavior" if asserts are disabled for performance.

And if in your use case you don't need the performance, you could just not enable those warnings and let it silently add all the checks it wants, and you'll have guaranteed safe (even if slow) code. The obvious concern of "but zeroes will appear unexpectedly!" can then be addressed by getting into a habit of always assigning zero the meaning of "not special; default; disabled; resource not available; operation not known to be successful" or similar, and then any unexpected zero that shows up will just neatly send your program down its fallback or error path rather than doing anything dangerous.

Division by zero, and other UB

Posted Sep 28, 2026 18:41 UTC (Mon) by ballombe (subscriber, #9523) [ Link ] (7 responses)

Integer division by 0 should trigger SIGFPE like with double.

Division by zero, and other UB

Posted Sep 28, 2026 19:12 UTC (Mon) by ballombe (subscriber, #9523) [ Link ]

Apparently x86_64 triggers a SIGFPE but aarch64 returns 0.

35 years ago the C standard committee should have used its weight to encourage CPU with standardized behaviour, but instead they tried to beat fortran at its own game and completely failed.

Division by zero, and other UB

Posted Sep 28, 2026 19:15 UTC (Mon) by Cyberax (✭ supporter ✭, #52523) [ Link ] (5 responses)

I think this is optional? Floating-point operations usually signal error conditions inline, using NaN values.

Division by zero, and other UB

Posted Sep 28, 2026 20:19 UTC (Mon) by jengelh (subscriber, #33263) [ Link ] (1 responses)

Floating-point operations usually signal error conditions inline, using NaN values

If only. Thanks to , it's not guaranteed.

Division by zero, and other UB

Posted Sep 28, 2026 21:30 UTC (Mon) by NYKevin (subscriber, #129325) [ Link ]

Unfortunately, we cannot even blame C for this misfeature. is just a portable interface to hardware-level behavior that has been burned into practically every microprocessor or FPU manufactured in the last N decades (for some value of N which people may debate in the replies, if they have nothing better to do). This also means that non-C programming languages either have to put up with it, or else grow a runtime that explicitly reconfigures it to some sensible default behavior (and then you have to deal with all the usual FFI etc. bugaboos implied by such a runtime).

Division by zero, and other UB

Posted Sep 29, 2026 8:52 UTC (Tue) by malmedal (subscriber, #56172) [ Link ] (2 responses)

IEE754 mandates that it be configurable. You can chose either a trap or inline signalling. If you don't the defaults vary.

Division by zero, and other UB

Posted Sep 29, 2026 9:32 UTC (Tue) by pm215 (subscriber, #98099) [ Link ] (1 responses)

I'm pretty sure IEEE754 doesn't mandate being able to configure trapping on any of the various floating point exception cases, because a lot of CPU implementations don't implement trapping. Notably, most Arm cores don't implement trapping on fp exceptions: architecturally it is an IMPDEF choice to support traps, and the only implementations I know of that do so are Apple's ones.

The IEEE spec says the default is "set the flag and continue execution", but it also muddies the other waters by devolving various aspects of this to the individual programming language specs.

Division by zero, and other UB

Posted Sep 30, 2026 6:17 UTC (Wed) by malmedal (subscriber, #56172) [ Link ]

It's a should, not a must, but it's specified in section 8.

Division by zero, and other UB

Posted Sep 29, 2026 8:32 UTC (Tue) by alx.manpages (subscriber, #145117) [ Link ] (1 responses)

Having defined behavior would benefit few use cases; and it would give erroneous answers that will hide bugs. It's preferable to entirely disallow division by zero when the operands are integer constant expressions, and UB at run-time, which analyzers can catch.

Division by zero, and other UB

Posted Sep 29, 2026 9:38 UTC (Tue) by ojeda (subscriber, #143370) [ Link ]

It's preferable to entirely disallow division by zero when the operands are integer constant expressions, and UB at run-time, which analyzers can catch.

For something like C, there is no need to use UB for analyzers to catch that -- they can (and already do) add checks like that just fine in practice, and if one wants to allow for that in the standard, one can introduce EB instead (if one is OK with a performance cost due to whatever defined behavior is chosen, of course).

Either way, what is best is to provide the user with the ability to be as explicit as possible: if they need UB for a particular reason, let them ask for it explicitly; if they know zero shouldn't happen but they don't need the performance in that particular spot, then let them use EB; if they need particular handling on the zero case, then let them ergonomically do that; and so on.

That is what also allows to easily read and understand (for both humans and tooling) what programs are supposed to do and what is unintentional.

Division by zero, and other UB

Posted Sep 29, 2026 11:26 UTC (Tue) by MortenSickel (subscriber, #3238) [ Link ]

This looks to me as taping a piaraya onto your boomerang. - it will sooner or later come back to bite you. (*)

I cannot recall one single case when I have written something where a x/0 returning 0 would make sense. In some cases I have a denominator I know can be 0, then I have to check for it and handle the situation differently if it is 0 or not. In other cases, I may have a denominator that should never be 0, and if I feel brave enough, I do not test, or I may test just in case.

What I definately do not want, is an accidential division by 0 returning something that afterwards looks like a reasonable answer, using that for some further calculations or checks and end up with a routine that seemed to run well but returns a nonsense value.

(*) Quote from last Edinburgh fringe festival.

Division by zero, and other UB

Posted Sep 29, 2026 16:07 UTC (Tue) by wtarreau (subscriber, #51152) [ Link ]

Note that if the divide by zero is often called "divide overflow", it's not without a reason. While you could play tricks with x/0, it's harder with the real, less known, overflow, which is to divide the max negative number by -1, for example -2147483648 / -1. This one turns to positive 2147483648 which cannot be represented as a positive number since it's the same value as -2147483648, hence the overflow. Same for 64 bits of course.

Division by zero, and other UB

Posted Sep 30, 2026 15:02 UTC (Wed) by scott (subscriber, #581) [ Link ]

Fun fact: there is actually a mathematical structure for a/0 = 0. They're called meadows, because they're nice fields (get it?). See for example https://arxiv.org/abs/1406.6878

Good news

Posted Sep 29, 2026 1:56 UTC (Tue) by marcH (subscriber, #57642) [ Link ] (1 responses)

Overall, he concluded, C is still a living language and is still improving. The C23 standard removed a number of problematic features, ...

It's easy to debate what programming languages everyone should use. Whereas this sort of work is barely visible and much more important - thank you! I don't want to choose between Rust and a safer C/C++: everyone should want both. They are not mutually exclusive and the more "security competitions", the merrier. Computers have been insecure and crashing for way too long and everything helps. Society will depend on C and C++ for at least a couple more generations.

Good news - Subset of a superset

Posted Oct 7, 2026 7:46 UTC (Wed) by swilmet (subscriber, #98424) [ Link ]

The C++ approach to make the language safer over time is to add features to it (so creating a superset of the language), and then to restrict the use of unsafe parts through compiler flags (taking only a subset of modern C++).

https://herbsutter.com/2024/03/11/safety-in-context/

Nothing prevents from doing the same for the C language. But C evolves quite slowly compared to C++.

Doubtful

Posted Sep 29, 2026 7:47 UTC (Tue) by taladar (subscriber, #68407) [ Link ] (20 responses)

I have my doubts with all of these efforts to fundamentally move existing languages very far away from what they are today and have been for most of their life.

The main benefit from using an older language like C (if there is any) is that the name identifies a certain language behaviour that existing C programmers know and existing code bases expect.

So while it is of course possible to change all kinds of things and still call the result C it is questionable at best to do so if it costs you the compatibility with the existing code bases and the existing programmer's knowledge.

At the same time though, it is orders of magnitude harder to ever get anywhere close to the benefits of a modern language like Rust that had the benefit of being able to start with a clean slate without having to consider backwards compatibility with existing code bases.

Even if you could add enough features to the language that new code had most of the benefits, you would still have to live with having old code in your process (unless you throw away backwards compatibility completely in which case, why not call it something else) and that comes with a myriad of issues when enforcing any sort of safety or security guarantee.

Doubtful

Posted Sep 29, 2026 8:41 UTC (Tue) by alx.manpages (subscriber, #145117) [ Link ] (12 responses)

I have my doubts with all of these efforts to fundamentally move existing languages very far away from what they are today and have been for most of their life.

I don't think it changes fundamentally. It's still the same thing, just changing some UB for compiler errors, and some other UB for defined behavior (the former is preferable).

The main benefit from using an older language like C (if there is any) is that the name identifies a certain language behaviour that existing C programmers know and existing code bases expect.

Old code that had defined behavior still works, and remains having the same meaning. The worst code using deep UB, it will probably now result in compiler errors; that's fine, since it never really worked. It can be fixed (and it can also be compiled under -std=c89 if UB is indeed wanted).

At the same time though, it is orders of magnitude harder to ever get anywhere close to the benefits of a modern language like Rust that had the benefit of being able to start with a clean slate without having to consider backwards compatibility with existing code bases.

According to Ojeda in Kernel Recipes, Rust already has applied breaking changes that silently change the behavior of some code that was valid and remains valid. That's something C has never done (IIRC) until very recently, where auto was changed to mean __auto_type. And it could be done with auto because no-one had seriously used it (there may be a few exceptions, but auto as automatic-storage duration was essentially unused).

So, even with many decades of advantage, Rust has had to do these things it didn't need to.

Doubtful

Posted Sep 29, 2026 9:32 UTC (Tue) by ojeda (subscriber, #143370) [ Link ] (11 responses)

According to Ojeda in Kernel Recipes, Rust already has applied breaking changes that silently change the behavior of some code that was valid and remains valid.

To clarify: when migrating from one edition to another (particularly Rust 2021 to Rust 2024), i.e. it is an explicit change, not a silent one.

The silent part was about backporting in Linux: if we tag a patch in Linux for backport, and the editions involved were to be such a pair, then the change would indeed be silent, because at the moment the process works in a way that makes it very likely nobody will notice the semantic change.

And thus why I would like to have a tool that the stable kernel team can run to identify such changes so that patches can be flagged, similarly to how they are flagged when conflicts happen.

I hope that clarifies.

Doubtful

Posted Sep 29, 2026 10:37 UTC (Tue) by alx.manpages (subscriber, #145117) [ Link ] (10 responses)

To clarify: when migrating from one edition to another (particularly Rust 2021 to Rust 2024),

Yup.

i.e. it is an explicit change, not a silent one.

That's more or less like changing the meaning of code from C17 to C23 (which has happened, with auto). In C, we'd call that a quiet change, because the one changing the language edition might not be aware of the code it is migrating, and thus of the implications of the change (essentially, the problem you face in the kernel with stable backports).

Here's what the current C Charter says about silent changes:

Avoid quiet changes

Changes that alter the meaning of existing code cause problems. Breaking changes that require diagnostic messages are easily detected. Avoid silent changes that cause a working program to behave differently without requiring a diagnostic message. Where this principle is violated, informative notes should be added to the Standard.

And here's what the C23 Charter said:

  1. Avoid "quiet changes." Any change to widespread practice altering the meaning of existing code causes problems. Changes that cause code to be so ill-formed as to require diagnostic messages are at least easy to detect. As much as seemed possible, consistent with its other goals, the Committee has avoided changes that quietly alter one valid program to another with different semantics, that cause a working program to work differently without notice. In important places where this principle is violated, the Rationale points out a QUIET CHANGE.

Which we've respected quite much, precisely because backporting issues are very problematic.

The exception is below:

auto x = 0.0;

The meaning of the above in C89 was int x = 0.0;, and in C23 it means double x = 0.0;. Hopefully, nobody will write code like that and backport it; at least that's what the committee hoped. I'm dubious of this precise change, precisely because it adds some unnecessary risk. On the other hand, there's the risk reduction in that C++ code ported to C would now mean the same, so maybe it was a good change. Since that code is weird in the first place, and avoided by C programmers in general, I'm not too worried.

I don't see any 'QUIET CHANGE' notes in C23, which I'll report as a bug in C23.

Doubtful

Posted Sep 29, 2026 12:01 UTC (Tue) by khim (subscriber, #9252) [ Link ] (3 responses)

You are conflating different C standards to Rust editions which is fundamentally wrong.

In C, we'd call that a quiet change, because the one changing the language edition might not be aware of the code it is migrating, and thus of the implications of the change (essentially, the problem you face in the kernel with stable backports).

Precisely — but that's because new features introduced in new C standards are not accessible for the code that's compiled for old C standard. You couldn't compile your code as C17 code and use _BitInt types, e.g.

Also: C offer no way to compile some library as C17 code (including things like macros) and yet use it in C23 code — while Rust remembers which edition was in use when macro was defined to permit precisely that mix. It's expected and perfectly normal to combine code for different Rust editions in one program.

This means that you may want to upgrade to later C standard to get useful features and then hit these "unexpected changes".

But in Rust that's not the case: you have access to all new features in all editions (except when new feature requires new syntax not supported by old edition). You only upgrade to a new edition specifically to opt in into these "silent behavior changes".

It's a bit silly to call behavior change "silent" when it's well documented and, more importantly, the whole reason new edition even exist!

Doubtful

Posted Sep 29, 2026 14:24 UTC (Tue) by alx.manpages (subscriber, #145117) [ Link ] (2 responses)

Precisely — but that's because new features introduced in new C standards are not accessible for the code that's compiled for old C standard. You couldn't compile your code as C17 code and use _BitInt types, e.g.

Compilers often implement new features that are backwards compatible, even in standard mode. So, actually, you can, and it seems very similar to what you say of Rust.

alx@debian:~/tmp$ cat bi.c
int
main(void)
{
 _BitInt(8) i = 42;

 return 42;
}
alx@debian:~/tmp$ gcc -std=c89 bi.c
alx@debian:~/tmp$ ./a.out; echo $?
42

Also: C offer no way to compile some library as C17 code (including things like macros) and yet use it in C23 code — while Rust remembers which edition was in use when macro was defined to permit precisely that mix. It's expected and perfectly normal to combine code for different Rust editions in one program.

Some experimental compilers do provide such a feature (IIRC). It's not something common or even desirable, but it exists.

You only upgrade to a new edition specifically to opt in into these "silent behavior changes".

Hummmm, sounds like a huge problem in 2099, when there might be dozens of editions with slightly different behavior. If one doesn't update code unless a quiet change is needed, then most code might stay on old editions, and a few lines of code might use wildly different editions.

Doubtful

Posted Sep 29, 2026 23:13 UTC (Tue) by mathstuf (subscriber, #69389) [ Link ]

Hummmm, sounds like a huge problem in 2099, when there might be dozens of editions with slightly different behavior. If one doesn't update code unless a quiet change is needed, then most code might stay on old editions, and a few lines of code might use wildly different editions.

There's no issue with that. It's not like Rust 2015 is going anywhere. It might be easier to think of newer editions allowing new _spellin

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →

Earlier on PAN

More in Technology →