On 7/20/26 02:19, David Hildenbrand (Arm) wrote:

We've the virto-mem config struct layout and the kernel source, so for
obvious fixes like a NULL check, static analysis is better than fuzzing.
Claude took a few mins to find me two examples:

Patch 1: virtio-mem: reject non-power-of-two device_block_size
This one is for virtio_mem_init() to check if
!is_power_of_2(vm->device_block_size)

Patch 2: virto-mem: validate region_size and usable_region_size
THis one checks region_size != 0 and vm->usable_reion_size >
vm->region_size.

An endless factory of "silly" checks like these are low hanging fruit.
"silly" is the right word.
"silly" in what way?

As in producing "silly" low-hanging fruit patches that don't move the needle
when it comes to security.

Seriously, I'm trying to figure out what you all care about here and
what exactly the threat model you want this driver to work in, and I'm
getting conflicting answers.

Either you all do worry about the "device" sending bad data and want to
protect from that, or you don't and you trust it.  Pick one please so
that we know how to deal with these bug reports we are getting.

For example, for USB we have said our threat model is:

   - we do NOT trust the device before a driver is bound to the device,
     so if a malicious device can do something to the kernel, the kernel
     needs to be fixed.
   - During the probe() call for a USB driver, the driver does NOT trust
     the device, and again, anything a malicious device can do to the
     kernel, the kernel should fix.
   - After probe() for a USB driver succeeds, it's up to the driver if it
     wants to validate all data coming from the device or not.  Right
     now, in general, the kernel trusts the device at that point in time
     so additional checks are discretionary and at the whim of the
     maintainer.

For that last point, I will note that some BIG users of Linux (i.e.
billions of Android devices) still explicitly do NOT want to trust the
USB device at this point in time, and are relying on the kernel to
protect the system from bad devices.  In that case, various patches have
been taken to different drivers and subsystems to play whack-a-mole on
while Android gets their act together to finally come up with a solid
defensive plan (like ChromeOS has had for a decade.)  It will be seen
which happens first, all drivers are properly fuzzed and fixed up, or
Android gets their act together and finally fixes their b0rked system
trust model.  I think Android management is relying on the kernel
community to do the kernel work as they keep refusing to staff the
userspace work that they need to do here...
Right, and for virtio devices trusting the device after probe is just extremely
questionable.

What changes during probe that the device suddenly sends us good data?

Why would a hypervisor that tried to break us before probe not try to break us
after probe?

And yes, I really need to write this up in a more solid document for USB
and get it into the tree, but at least this email thread has forced me
to write down the above :)


So, again, for virtio drivers, what exactly do you all want to say is
your threat model that the drivers need to handle?  Can you all agree on
something please?  Otherwise, for new developers like Hari, this is
totaly confusion as to what they should be doing.
Well, I am also totally confused why we end up checking against some MUST
clauses in the spec, but not against others.

I am very much in favor of making virtio-mem completely safe to use even in
coco, where it is currently not used at all.

If it's really about "don't let a device trigger any unexpected kernel code
execution by sanitizing all input data", fine with me. We should do exactly
that. Try checking all MUST clauses etc.

But I don't think doing the "low hanging fruit" adds any security. It should be
done properly or not at all.


I think you're mistaken in assuming that these "silly" fixes don't move the
needle at all when it comes to security. We can't really quantify these
things, and one extra NULL check by itself is probably not going to make a
meaningful difference. But, IMHO, this should be treated as the opposite of
"death by a thousand cuts", many small hardening improvements, each
individually insignificant, can collectively make the kernel substantially
more robust over long periods of time.


Thanks,

Carlos


Reply via email to