I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.
Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.
Not sure I will give up on smolvm though.
I have scripts to launch an instance per task. Nono wraps my coding agent.
I provide the git clone for the specific task.
It works really well.
I'd also add that MXC is still sharing the host kernel.
i.e. on mac MXC uses seatbelt, so the agent shares your kernel.
but smolvm gives each task its own VM with its own kernel, same on mac, linux and windows. + additional layers like preconfigured seccomp, landlock, wherever it makes sense
The "audit" and "debug" mode are specially useful. I use bubblewrap and using a new harness with it is usually a couple of rounds of wack-a-mole with strace to figure out all the harnesses dependencies.
They have this for Windows and Linux, but it's sadly missing for macOS - see the support table here: https://github.com/microsoft/mxc/blob/main/docs/backends/sea...
Things macOS is missing include "Allow/deny by hostname" and "Allow/deny by IP, CIDR, port, or protocol".
The rest all looks great, and if you are on Linux or Windows those restrictions don't apply.
I guess this is the universal challenge of building an abstraction layer over multiple different technologies.
smol machines actually support exactly those things across macs,linux, windows btw: https://github.com/smol-machines/smolvm/blob/main/AGENTS.md#...
here's a snippet of how it looks like to configure that:
[network]
allow_hosts = ["api.github.com"] # hostname, also allows its subdomains
allow_host_patterns = ["example.com", "*.npmjs.org"] # exact names, or *. for subdomains only
allow_cidrs = ["10.0.0.0/8", "1.1.1.1"] # IP ranges or single IPs
[[network.credentials]]
name = "github"
environment_variable = "GITHUB_TOKEN"proxy in the middle (but cert pinning problems)
or DNS filtering? (but agent could have "memorized" stable IP)
Memorized IP: doesn't work, the vm can only connect to an IP if it came from a DNS lookup of an allowed name. Any other IP is blocked.
A bit of "shared responsibility" philosophy kicking through but I try to have good defaults
Of course, that only limits HTTP; and not other forms of network requests.
https://github.com/anthropics/sandbox-runtime/tree/main#as-a...
const config: SandboxRuntimeConfig = {
network: {
allowedDomains: ['example.com', 'api.github.com'],
deniedDomains: [],
},
filesystem: {
denyRead: ['~/.ssh'],
allowWrite: ['.', '/tmp'],
denyWrite: ['.env'],
},
}I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).
Does anything close to this exist yet?
Don't want these capabilities? Do nothing, that's the default state.
However outside mainframes and micro-computers, or niche deployments, adoption has been a challenge.
But isn't this all a bit adhoc? The fundamental presumption is of unrestricted capability and then the OS provides options for restriction.
Perhaps I've misunderstood it but I think capabilities are inside out from all of that: You need to walk, so here are some legs. You need to swim so here are some fins. As-is it's more like: you can't go over there so here's an ankle monitor.
I'd be way more confident building upon something I can grasp and understand. [2]
[1]: https://ghloc.dev/microsoft/mxc [2]: https://github.com/sandbox-utils/sandbox-run
https://ghloc.dev/microsoft/mxc?branch=main&locsPath=%5B%22s...
That site thinks this file has 2.9k sloc and doesn't seem to parse rust comments. In reality, there's only 1,465 sloc; and 635 loc of tests.
Definitely nowhere near 350k sloc.
--
As for your sandbox run: it's a single-contributor project, seems to have only have basic smoke tests, and has a few major/critical security issues:
* _generate_seccomp_filter compares newline-deliminated syscalls, against a multi-line blocklist, meaning the entire function doesn't block anything and is essentially a no-op.
* Main script invokes working directory's .env as shellcode, before switching into restricted filesystems and dropping capabilities. Attacker-controlled .env can run shellcode with full privileges.
* Lots of race conditions which I haven't verified, but doesn't really matter.
I'd make PRs, but I don't think it's a good idea to try and DIY a sandboxing system in bash with minimal SLOC as the target in the first place. I'm also slightly concerned that most of your comments on HN seem to be promoting this repo?
Thanks, I see there's a slight (~50%?) overestimation there, but then again, even unit tests and comments in a target programming language count as syntactically correct code that needs to be evaluated and reasoned upon. I'm not that familiar with Rust's runtime introspection features, but in languages like Python, even the comments can directly affect code (e.g. `Foo.__doc__ = Bar.__doc__ + SOME_ANNEX`).
> _generate_seccomp_filter ... the entire function doesn't block anything
Many thanks! I've applied a fix—it's a single line added. The missing test is pending a runnable that invokes one of the forbidden syscalls. As I have no qualms about force-pushing around a repo that nobody forks, happy to credit you(r LLM) proper!
> Main script invokes working directory's .env as shellcode
The sandboxed process can't overwrite existing .env files [1], but it could create a new $PWD/.env file, hoping to "escape" at next sandbox execution. That's a valid concern I'll have to think about some more.
[1]: https://github.com/sandbox-utils/sandbox-run/blob/c97d065184...
> Lots of race conditions
I sometimes experience "Slirp not ready in time" [2], but it's due to a so far unexplained upstream issue [3]. I you have time/tokens to spare, I'd appreciate those PRs and further similar feedback!
[2]: https://github.com/sandbox-utils/sandbox-run/blob/c97d065184... [3]: https://github.com/rootless-containers/slirp4netns/issues/35...
Don't know whether it's a good idea. It sure has got its issues. But even as the SLOC count and the number of bugs metrics are proved correlated in literature [4], min SLOC is not the primary target—a reasonably graspable and stable composition of few dependencies is. Whereas overreliance on third parties nowadays often ends with a rug pull one way or another. We simply can't count on this "MXC" (...) to be maintainable/non-archived even a year from now, just when I'd get it all properly integrated and set up.
[4]: https://softwareengineering.stackexchange.com/questions/1856...
> slightly concerned
Oh, I certainly wouldn't like to limit myself to promoting just this repo! ^D^ HN is a good venue, lots of smart people around! I see everyone shilling their own sh** all the time. Often in green usernames. :shrug:
You would not be aware of the amount of trust you are putting into that.
That being said, a month of MXC has more line changes than 2-3 years of runc
Even the bubblewrap integration docs are basically a stream of consciousness vibe splat. I have approximately zero confidence in the results.
The API for their Rust mxc-sdk looks nice but
- their "sdk" has binaries and the build script has logic for them
- their build scripts do windows-exclusive work on all platforms
- not putting some of the backends behind features causes more build script work (and that work will break on future Cargo versions)
- at least some of the remaíning build script work doesn't need to be a build script
- it seems pretty dependency heavy
https://wasmer.io/ (some cool examples here https://wasmer.sh/)
I tend to use Wasmtime, V8, WAMR, and browsers. The point of WebAssembly is that there are many runtimes.
wasip3 is not stable yet, but it has a lot of nice changes (compared to wasip2) for integrating with async code
Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.
I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.
I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.
"Im really fkin worried we're all building the same thing"
Everyone has been building a harness/sandbox the last 6 months. Ive seen dozens and dozens shared in discords.
Even companies are totally stuck focused on the same paradigms.
The previous iteration of this was RAG/Chat interfaces. See PewDiePie's project. Last month it was briefly everyone building the same classifier.
Peter Thiel, gave a lecture about this same phenomenon 15 years ago likely because he observed the same things going on during other hype cycles. Everyone building the same things. Its called something like "Dont build the obvious thing"
This is why Im moving towards hardware for personal projects, it forces me to be much more creative and think outside the "How can I make something AI adjacent/powered" trap thats so easy to fall into in pure software right now.
e.g. this one puts multi-platform support as a high requirement, a requirement that OpenShell doesn't fulfil (and likely won't given it's architecture/goals).
History is littered with tons of super cool OS features that didn't manage to gather enough market share and ended up as cool futures, and fodder for 'we invented the future 20 years ago' style articles.
In reality, however, knowing how much difference there is between OSes, how tricky it is to configure these things to make them actually useful, and how bad Microsoft products are, I'm not enthusiastic about this project -- there are so many others on the market already, and I'll wait to see if this gains traction.
(Notice that on MacOS it only supports seatbelt? That's not nearly the same as microvm.)
But as always, there are rivals trying to set their tone on what’s the standard. We all wish there was one unified agreed concept that will work but I guess the most common one will eventually survive.
Just as Microsoft in a sense embraces Linux with WSL and also Apple has their virtualization framework.
I hope we’ll eventually get unified model management system to include also permissions designed properly
I want to support sandboxing for my app, and commands it launches, but there is nothing that actually works across all the environments I want code to run in. So yes, if I had one tool that could be an abstract interface and let the user set up and configure their sandboxing completely separate to my app, and do that at run time based on declarative policy - it would be handy.
I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far
Why we need this layer of complexity? Or its mainly for big company that need control ?
I’d personally opt for SELinux in such cases
https://learn.microsoft.com/en-us/windows/win32/secauthz/cre...
Funny, now that it's documented, the API name will be stuck with this name forever.
Microsoft is already a Chromium contributor so it's not like they lack in-house expertise or something.
https://wiki.mozilla.org/Security/Sandbox
https://chromium.googlesource.com/chromium/src/+/HEAD/docs/d...
But seriously, we need docker for models like years ago. I dont want these things running with the ability to run rm -rf /
its should be treated no different than wget | sh
Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.
If that’s too much of an ask, at least reference the code you found for things like “mxc sandbox escape” or “bubblewrap setuid is wrong”. Those claims require evidence.
"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."