Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Overall, this seems like one of the weaker Mill talks. Since they apparently don't yet have a real OS running real software in simulation, they probably haven't had the ability to test the ideas that affect how software is structured at a higher level.

They don't provide nearly enough ways to transitively grant permissions. Using the mechanisms discussed in the talk, it doesn't seem like you can implement a simple asynchronous queue of units of work to perform, each having their own permissions. The belt architecture encourages these sorts of second-class mechanisms that have to be used in a rigid way, because the details can be hidden in the belt and not be exposed architecturally.

Unless there's something else not mentioned in the talk, it seems like you still need to trust the OS, because when the OS is asked to allocate a page for a spillet there is nothing stopping it from creating a virtual alias of that page elsewhere and allowing another thread to read its data.

The mechanism to support fork() is a total kludge. Why have a single address space if you're just going to add segmentation in such an ad-hoc way for a single use case? Just run the original binary in emulation until exec() or something like that.



> The mechanism to support fork() is a total kludge.

Fork is a kludge. It also happened to be easy to implement on the hardware available at the time and we've been stuck with it ever since. So I happily forgive the mill team that their fork() implementation looks like it is a total kludge, it would be highly surprising if it were not.


>Unless there's something else not mentioned in the talk, it seems like you still need to trust the OS, because when the OS is asked to allocate a page for a spillet there is nothing stopping it from creating a virtual alias of that page elsewhere and allowing another thread to read its data.

The Mill is hardware and architecture, not policy. If you want to use such an OS then you are free to do so. The Mill is designed to efficiently support micro-kernel OSs. Note: micro-kernel, not no-kernel. There always will be a Resource Service that owns the machine. It will be a couple hundred LOC, small enough to be correct by eyeball or proof. Contrast your choice of monolith.

The OS is not involved in allocating spillets. Spillet space is a large statically-allocated matrix in the address space. It is not allocated in memory, only in the address space. As soon as you allocate a turf id and a thread id you have implicitly allocated a spillet. Only on spillet overflow is allocation necessary. Whether allocating turf or thread ids requires OS involvement depends on the policies and models chosen by the OS designer.

When first created the spillet data lives only in backless cache - no memory is allocated. Only if the spillet lives long enough to get evicted from cache is actual memory allocated, using the Backless Memory mechanism described in our Memory talk. The root spillets of apps will live that long; transient spillets from portal calls will likely live only in cache. Consequently truly secire IPC/RPC using Mill portals has overhead, both app and system combined, of the same magnitude of an ordinary function call.

> They don't provide nearly enough ways to transitively grant permissions. Using the mechanisms discussed in the talk, it doesn't seem like you can implement a simple asynchronous queue of units of work to perform, each having their own permissions.

There is a "session" notion that addresses such things. Unfortunately the talks are far enough into details that they must contain background and introduction slides for the viewers who have not already done (and retailed) all the other talks. This limits the amount of new material that can be covered in a single talk, and sessions didn't make the cut this time. We'll get to them.

> The mechanism to support fork() is a total kludge.

Agreed; there seems to be a Law of Conservation of Kludgery. We had as a minimum requirement that the architecture must support Unix shell. The only real problem is fork(). Would that we could issue an edict banning it.


They've repeatedly said they want to run current software well, presumably including software that forks and does not follow up with exec (regardless of how ill-advised that may be; obviously opinions vary on this). Trapping to emulation on fork would seem to fly in the face of this.

That being said I'm still not sure how you can get all of fork's semantics out of this mechanism...

(For that matter various sneaky VM aliasing tricks ["magic" circular buffers mapped twice in a row so that any size and position block is contiguous, same file mapped in multiple processes with different (non-contiguous) memory mappings] are going to fail miserably with the virtual-addressed cache. It may well be worth not being able to do those things for the benefit of moving the TLBs, but it also flies in the face of "will run current software.")


(team Mill)

> "magic" circular buffers mapped twice in a row so that any size and position block is contiguous

Yes this exact case is one of those that won't work. There are a handful of other things that don't work, like people who roll their own coroutines using assembler. We promise 'don't rewrite, just recompile' but your codebase has to be 64-bit clean and not use any assembly language for any other ISA and not make other unportable assumptions about the hardware.

I can't think of any software that actually uses these circular buffers that wouldn't be worth adapting to the Mill rather than adapting the Mill to them; can you? ;)


How do you plan on supporting user/process level coroutines and multiple stacks / stack switching within the same turf? Obviously the CPU needs to be involved because you've got all the ancillary belt and spiller state that needs to be switched as well.

For that matter have you guys touched on how stack unwinding is going to be handled yet? How do I implement exception handling, longjmp, CL's non-local GOTO, etc?


The next talk will be on threading, and will address all this. The IPC and threads talks belong together, but are too big to combine in one talk unfortunately.


the fork exposition was weak. admittedly fork() was a mistake and constrains alot of implementations in strange ways. i still dont understand how local/global exactly matches the semantics of cow.

transitive permissions are capabilities.

while i'm sympathetic to the lack of market appeal to a capability based system, doesn't it seem like you could implement posix on top of one by compromising it? fd transfer over unix domain is already halfway there.

seems like a better alternative.


The biggest problem with caps and legacy apps is not the semantics - a caps systems can emulate POSIX with no problem. The problem is the data representation: you can't fit a capability into a pointer, so all the data layouts change. Sadly, there are tons of C programs that make rash assumptions about data layout, and they would all break.

The guys are Cambridge have running caps systems that store the extra info in outboard data structures. We judge that the overhead is too great for commercial success. Customers buy benchmarks, and there are no security benchmarks.

'Tis true 'tis, 'tis pity. 'Tis pity 'tis, 'tis true.


> i still dont understand how local/global exactly matches the semantics of cow.

It's not, it's "copy on reference".


No, it's aliasing, just as any fork() is. COW lowers the cost, for Mill as for any other. The Mill fork duplicates the address space, not the memory. The memory is duplicated page by page on demand, i.e. copy-on-write. Mill paging is quite conventional; it's the address space (SAS) that is different.


They were light on the details, but what stops an OS from mapping the same physical page into two distinct local regions of the address space and implementing copy-on-write as usual?


Presumably nothing, but it doesn't seem like that would break anything?


(team Mill)

> They don't provide nearly enough ways to transitively grant permissions.

Portals are synchronous and transient permissions can only be used for the duration of the call and by the same thread. Asynchronous isn't so 'simple' because its about lifetime. With the synchronous portal the caller knows that the callee cannot retain any access to the buffers that were passed, and can reuse them safely. If those buffers were put on an asynchronous queue, when would the memory be safely reused and when would the owner know that? If you want asynchronous queues, you either have to have a buffered model like Unix pipes etc or you have to have a global hardware-implemented GC that somehow spans turfs and becomes part of the Trusted Computing Base (TCB) (shudder).

> it seems like you still need to trust the OS, because when the OS is asked to allocate a page for a spillet

Well there's plenty not mentioned in the talk and you touch on one aspect :) When a spillet overflows the extension cannot be in the reserved space, so space must be 'carved out' of the part of the address space where programs also have their needs carved out. Someone has to do the carving, whether its for spillets or for programs, and that someone has to be trusted. The problem isn't aliasing (we're Single Address Space), its that they can simply give as many permissions to it to as many turfs as they choose. There is a possibility that the carver is in the BIOS, but any which way there has to be a turf that can do this. This turf is obviously part of the TCB.

> Why have a single address space if you're just going to add segmentation in such an ad-hoc way for a single use case?

Unfortunately there isn't much market for general purpose CPUs if they are fundamentally unable to run Linux ;)

As you will almost certainly be running a Unix, and as you almost certainly will be using libraries that may fork, then all your normal heap and data stack pointers are going to be Local. Shared mmap pointers will be Global, as will your code. We can hope there is a flag this Linux lets you set that says that you forego the ability to fork() and in return all your pointers can be Global, because the Local puts some constraints on your address space use which become clearer if I explain them:

The local bit works like this:

If a pointer has the bit set, then before use it is mangled with a special register called the Local Space register. The hardware does this every time it uses a pointer. Each turf has its own Local Space value, and you can think of it as a simple offset, so if the address is 0x1...10 and the local space is 2 then the effective address in the global space is 0x1...100.

Now 64-bit adds are relatively slow for this use-case because we really want the effective address asap in the FU so a probable implementation of the Local Space is XORing it into the pointer instead of adding.

When a process forks the OS has to find a position in the global address space where the allocated ranges used by the program are free for the child and then set the new Local Space register for the child appropriately.

> Just run the original binary in emulation until exec() or something like that.

All modern OS don't actually COW until there's a page fault, and I'd expect them to use that trick on the Mill too. So the local bit makes it possible to fork(), but the hole-searching is lazy and only happens if you actually use it.


Hi, capability-security folk here. Have you considered adopting capability-security language for describing what's going on here? It seems to me that Mill portals are very much like capability-designed syscalls. In particular, I am reminded of the "no stale stack frame" rule from Monte [0], and its interesting justification:

"Since Monte permits mutable state, one author’s code’s behavior could be affected by another author’s code running further up the frame stack. Stale frames make comprehension of code much harder as a result."

I wish that we were having more conversations and classes like [1] to get knowledge of KeyKOS and other capability-safe kernels out to CPU makers. We have ideas to share!

[0] http://monte.readthedocs.io/en/latest/semantics.html#scope-i...

[1] https://pdos.csail.mit.edu/6.828/2009/lec/l-microkernel.html


The Mill grant-based model is semantically quite similar to capabilities, but it associates protection with the accessor (thread/turf) rather than the access (pointer/capability). This lets us preserve the size of a pointer, which no one knows how to do efficiently with capabilities.

The difference between the two models is visible when you pass a graph structure across a protection boundary. With caps is is easy to pass the whole graph, and hard to pass only one node. With grants it is vice versa.


Ah well Norm Hardy (KeyKOS) is on team Mill; the Mill is very much a bunch of caps vets :)

The Mill isn't a Capability-based Addressing machine, but it is very much a machine that hopes to push normal programs and operating systems in the Capability Object Model direction with first-class support.

Now you are using a bugmenot login so I have no idea who you are, but if you're a caps vet we'd love to know your thoughts :)


Curses, time to switch accounts.

My thoughts are mostly that I don't want to wait a decade for a Mill. Find me at SPLASH/OCAP if you want to chat.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: