upvote
I suppose there would have to be a capability based model in conjunction with a user oversight model and a time model.

https://en.wikipedia.org/wiki/Capability-based_security

Thus some agents with higher capabilities can only be run with user oversight at the same time.

Some agents can not be run during some part of the day - for example these agents can not run within two hours of office closing time, and cannot run on weekends.

Maybe also the idea of agents writing code - throwing "out fully-fledged programs that you have to approve or reject permissions for."

Would work better with a capabilities based model where you choose capabilities for the program before hand, meaning the capabilities are not written by agent itself, you read through the code, some of it looks hairy but everything is fine, but oh no dumb human missed the part where agent writes to system32! But luckily enough the program you were expecting actually needed no write capabilities and thus when it tries to go past its assigned capabilities that part of the program fails and the exception is registered.

Googling it seems like lots of people have thought this (at least where Capability based security is concerned), which seems reasonable to me as it also seems pretty self-evident it must be this way. Have not really seen anything about time based controls but then that is probably because I'm not devoting a lot of effort as I am just doing a bit of procrastination to build up the energy to finish something off.

reply
A language that revives capabilities, brings them up-to-date, and works in the modern environment is my #1 request from the programming language community right now. I don't need another language with sum types and higher-order functions and a functional focus. I need a language with capabilities. That language may have the other goodies as well, sure, no problem, but we all need capabilities.

I've done some stabby stabs at a design for it, using an AI as the rubber duck. My initial research indicates that the field of "static language that natively supports capabilities" is surprisingly uncovered and there may be a rich field there. E, the closest match, was tied at the hip to Java, which has some advantages but also comes with disadvantages for languages that are trying to do something as exotic as this. Other existing work was on dynamic languages, and hardly rose to the level of "practical for any use" let alone something that could solve our supply chain issues.

My issue is primarily that the reward for successfully designing a language and creating a community around it is that you're in charge of a language community... and, uh, my personality is not suited for that, that sounds more like something I'd pay to avoid then something I'd spend months and years of hard work to attain.

(My advice to anyone doing this is to spend some time with the AI researchers to find the existing work on the topic, not to just sit down and sketch out your initial ideas and run with them. Learn from the past. Expect this to be weeks and probably months of just thinking and noodling before you get to a design. Also, don't try to hook deeply to an existing language, as tempting as it is. This is way too large an impedance mismatch with existing languages. Any external code has to be treated like a nuclear bomb anyhow.)

reply
Been working on something like that for years: https://www.firefly-lang.org/
reply
if firefly has no nulls, how do you indicate that a value is unset?
reply
In Firefly (as in Rust) you can define fields as Optional, so you can do Option[String]; that lets you say "this variable is a String but it might not be here". That lets you then check to see if something is set, rather than checking to see if it's null.

In Rust an Option is a separate thing that you need to disambiguate to use. For example:

    match result {
        // The division was valid
        Some(x) => println!("Result: {x}"),
        // The division was invalid
        None    => println!("Cannot divide by 0"),
    }
Likewise in Rust, you can't have a null pointer, but you can have an Optional pointer, which is either a pointer to something or is not anything.

Firefly seems to have a similar case structure, though the first example I could find is in the Exceptions section: https://www.firefly-lang.org/reference/exceptions

    grabOption[T](option: Option[T]): T {
        | Some(v) => v
        | None => throw(GrabException())
    }
reply
oh duh! sorry, morning brain didn't think things through -- i do love Option types!
reply
What is a capability in terms of programming language design? It sounds more like the sort of thing that would belong at the standard library level, where builtin APIs are guarded by flags.

Deno has something vaguely built in with permissions flags, and old school Blackberry (at least in the J2ME days) had permissions settings for almost everything that an app could do, but again, those are all external to the language design itself.

reply
If you want a fun read that might educate via entertaining I would read Satan Comes to Dinner by Douglas Crockford https://www.crockford.com/ec/dining.html

I wrote a bunch more, but I have a habit to verbosity when capabilities come in which I should attempt to combat, and have deleted it. Crockford says these things better than I can anyhow.

reply
In this context, a capability is something that allows the code, or the transitive closure of the code that it may call, to access some particular function, to put it very briefly. So you could have a single function that, if accessed in one manner, is permitted to read from the directory /tmp/blahblah, but accessed in another manner, is permitted to read from the directory /home/zdragnar/.config/myprogram, and it is guaranteed by the language and runtime that the function will never do anything else on the file system. Or, even more importantly, it can be guaranteed that "from this code, nothing, no matter how the code is arranged, can access the file system at all".

This has massive overlap with a lot of things, like capabilities as implemented by Linux, effects systems, monadic data types as a not-really-very-good capabilities system (Haskellers have been playing with this for years and nobody really loves this approach, many practical problems beyond the scope of this message that would affect any language that tries that approach), dependently-typed programming, and so forth.

It is not a flag, though; flags can't handle that "transitive environment" aspect. It is also granular on the level of the programming language. This would allow you to do things like have your program be given access to a given part of the mobile file system using the mobile OS' permissions, but you could know beyond a shadow of a doubt that the image library you are using can not at any point access the file system, no matter what changes the author makes to it, because you can just look at the capabilities given to the image library and see that file system access is not among them. This is where real opportunity is over the next few years, in my opinion, because supply chain attacks are going to continue to get worse. A neat aspect of this approach is that it makes huge swathes of the ecosystem unattractive targets by statically ensuring that they can't sneak anything in to something that doesn't need file system or network access, so hackers won't even attack those libraries. Thus the ecosystem can concentrate on monitoring just the high-touch libraries that need to access high-risk resources.

(I should make it clear that the image parsing libraries can be passed a file; what I am saying is that they can't spontaneously originate arbitrary file system access in a system like this. Really what they would get is probably a "stream" and they would be forbidden from poking into the stream to see what it is made of, at which point, if some other code handed it a file presumably it meant to do that, but it does not give the image library any ability to do anything else with the filesystem.)

Moreover, if such a benefit was available, that would tend to have people squeeze down those dependencies as much as possible too, e.g., the aforementioned image library. You don't need file system access to parse images, that's just some convenience functions easily worked around that are provided because why not? The number of things that truly need direct high-risk access can actually be surprisingly small, and often, the application can also easily scope the permissions down quite tightly so the HTTP request library is limited in what it can hit, etc.

We actually have some semi-decent stabs at capabilities at the OS level; we can quibble with them but they are there. But inside an OS process, broadly speaking, anything can do anything in the vast majority of programming languages. The only way to be sure that the string concatenation function doesn't start crawling your file system looking for crypto keys is to examine the code, most languages have no ability to tell it that it can't. There are exceptions, like the aforementioned Haskell, that have at least some ability to do this, but this is an HN post, not a complete guide to a major topic. Really this is more about loading the reader up with keywords they can hit Google or an AI with.

The term is overloaded, too; Pony has something it calls "capabilities" but it really resembles more a sort of response to Rust's borrow checker, and if there is a way to lift it into this style of capabilities coherently it isn't clear to me. And even if you did, the entire rest of the ecosystem wouldn't support it, which is one of the reasons why this has to be a new language. You can't bodge this on to the side of an existing language.

(Plus, IMHO, there are some other ideas this may shake loose. Programming languages seem to be in a rut right now. My crack in my previous message about sum types and such isn't really about those things but the way almost every language going by is just a respelling of previous languages, churning over some other iteration of "The Perfect 2015 Language" that is already covered by any number of existing projects. I don't know that there's a lot of room there anymore. We need something big. Once you try something big the design will inevitably lead to other interesting things nobody else is trying either. Capabilities is one distinct possibility... like I said, if you dig in to the history you will discover there are entire huge segments of the capabilities space that haven't even been tried. If nothing else, if you are a PL nerd, I guarantee it'll be fun to explore those spaces that almost nobody has covered. No criticism intended to those who have, who have done a good job. It just hasn't been enough people and enough exploration to truly map the space.)

reply
> My initial research indicates that the field of "static language that natively supports capabilities" is surprisingly uncovered and there may be a rich field there.

You want to look for white papers that talk about object-capability systems. It's a fairly old and well-trod area of research. The E programming language[1] was all about that, and it was pretty late in the game on this stuff.

You emphasize natively, but the problem is that's not really well defined. For static capabilities, you're just essentially asking for a suffciently strong module system with parameterized abstract data types. It's literally a subset of the grammar and what it's designed to express. Mark Miller (one of the creators of E) demonstrated that[2].

The knock on effect of that quality is that anything which fulfills that requirement natively supports capabilities. It's part of the grammar. Doesn't even have to be object-oriented. A hackjob demonstration of an SML filesystem library with a brand/mint object capability pattern:

brand.sig:

    signature BRAND =
    sig
        type token
    end

mint.sig:

    signature MINT =
    sig
        include BRAND
        val mint : unit -> token
    end

makebrand.fun:

    functor MakeBrand () =
    struct
        type token = unit ref
        fun mint () = ref ()
    end

filesystem.sig:

    signature FILESYSTEM =
    sig
        type token
        val readFile  : token -> string -> string
        val writeFile : token -> string -> string -> unit
    end

filesystem.fun:

    functor FileSystem (B : BRAND) :> FILESYSTEM where type token = B.token =
    struct
        type token = B.token

        fun readFile (_ : token) (path : string) : string =
            "contents of " ^ path

        fun writeFile (_ : token) (path : string) (_ : string) : unit =
            ()
    end

trusted_fs_setup.sml:

    local
        structure FileAuthority :> MINT = MakeBrand ()
    in
        structure FS :> FILESYSTEM = FileSystem (FileAuthority)
        val rootFileToken : FS.token = FileAuthority.mint ()
    end

trusted_fs.cm:

    Library
        signature FILESYSTEM
        structure FS
        val rootFileToken
    is
        brand.sig
        mint.sig
        makebrand.fun
        filesystem.sig
        filesystem.fun
        trusted_fs_setup.sml

Now for any given library using the trusted_filsystem library:

  val doc = FS.readFile rootFileToken "/etc/motd" (* Works fine *)
 
Delegation is function application:

  fun helper (t : FS.token) = FS.readFile t "log.txt"
  val log = helper rootFileToken  
And these all fail:

  val fake : FS.token = ref () (* Trying to forge a token *)
  val t = FileAuthority.mint () (* Trying to bypass the trusted kernel in trusted_fs_setup.sml by calling the mint *)
  
  (* Trying to self-issue authority by making our own brand and mint *)
  structure MyCap = MakeBrand ()
  val t : FS.token = MyCap.mint ()   (* type mismatch *)

What's nice about this is... it's just normal modular programming. It's a very natural grain. It's also completely compile-time, no runtime overhead.

You can also do a lot of this with phantom types, and it'd be much more terse and easier to handle dynamic capabilities and stuff like a capability algebra, but it ends up way less auditable and is easy to have subtle errors which defeats the point. Also compiler errors will be much more opaque. IMO needing to manually make wrappers for composite capabilities, or to handle dynamic capabilities, is the lesser of two evils. With higher order modules, those problems go away entirely.

[1] - https://en.wikipedia.org/wiki/E_(programming_language)

[2] - https://homepages.ecs.vuw.ac.nz/~kjx/papers/ARND2018.pdf

reply
I think ultimately what it looks like it containing the blast radius if an agent does something bonkers.

The best case would be putting an agent in a VM and mounting the working directory there. Then you can allow it to run somewhat arbitrary actions while still being able to turn off the vm and restart it in a clean state.

The issue is, of course, that it doesn't fully prevent all possible problems an agent can cause. exfiltration is, IMO, basically impossible to stop. LLMs are exfiltration machines. The basic premise of all of them is "send us your code and a prompt and we'll do something good with it. But also if an agent decides run a command which installs a worm on a device on the network, you are hosed.

reply
The system I'm comfortable with is to set the agent up as an unprivileged unix user, with no ability to change system configuration and no access to any files I didn't specifically give it access to. Need to let it access a file or a directory? chmod is your friend.

Second, it can pull from git, or submit a pull request, but not directly push. We have an existing system of code review for that, now also augmented by llms.

Thirdly, prevent it from sending anything but get requests to anywhere you don't want it to post stuff, with firewall configuration.

After that, turn the horrible security theater of it asking permission for anything off. So far we have had no incidents. It could of course still pull a malicious package from somewhere, that exfiltrates code using GET, but at least it can't send any credentials or user data over.

reply
The way our Claude Codes are configured at work is pretty nice. There are directory patterns it can't access, like .local or .config, so when it needs to it creates a throwaway scratch pad and asks you to put stuff there; screenshots, text files, command output, etc.

I was having it diagnose a GNOME extension and had to get it a copy of the code to work on; it would then write out a Python script to do the patching (which I could inspect beforehand) and have me execute it.

Not having access to .local or .config can be irritating sometimes, but it's nice to know it's not just going to exfiltrate my docker or gcloud credentials.

reply
After watching a mid-tier offering chain together tools like it was a gorilla escaping the zoo I just gave the model its own box. I don’t have time to deal with that kind of nonsense.
reply
I cannot help with your actual but this is giving me mild ptsd flashbacks to everyone on hn/slashdot constantly repeating how simple and perfect unix security is, just use user accounts!

As if the most valuable thing on my pc was running a program on the gpu or the printer as opposed to my email account.

reply
No need to go back to the Slashdot days. Look no further than 2 days ago to find someone arguing the kernel is uniquely important (here for reliability): https://news.ycombinator.com/item?id=49181366
reply
On unix your email account is part of the filesystem.
reply
I don't think there's a way to make it secure while still permitting it unprompted external access

Eg: Any web request is a security vulnerability, there's no way to do it if the web requests are being made maliciously

Say that we have an agent with access to get requests, solely to a single site https://yoursite.com without subdomains. In this case multiple requests can be sent, and the time between requests can be used to exfiltrate personal data, similar to the coffee shop attack but without the subdomains. If the AI is able to make requests in any form, some information can be leaked, where the amount of leakable information is tied to information theory content of whatever side channel is being used. The only 0 information channel is.. never to make a request

You could also completely trust the 3rd party you're connecting to, but that to me seems like a hard error in the modern internet

reply
What a serious security model for a meatbag agent looks like? No, but seriously, an admin in a small org is a huge key-person risk in that they (or their stolen creds) can wipe enough and quick enough to effectively disable the business altogether.

More security conscious admins will at least segment their creds and implement four eyes principles somewhere, but were are back at square one of "asking user for confirmation".

Larger orgs, even if by necessity, segment their human agents, their creds and plaster four eyes principle liberally. But this relies on safeguards against agents colluding and ignoring some inputs, which sounds a bit scary for artificial agents.

Say you implement some swarm of agents, where access-enabled sub-agents are extremely restricted with system prompts and some access filtering. Then none of the agents in the swarm should be able to spawn themselves, otherwise a rogue agent can overwrite any safeguards. That, again, leaves the user with manually approving/denying network requests / hosts / sessions.

While I don't like anthropomorphising LLMs, the problem domain seems quite damn close to that of a key person going rogue within an org. The general solution seems to be liberal amounts of trust and ~~sweet compensation~~ gaslighting about replaceability.

reply
"What a serious security model for a meatbag agent looks like?"

Yes, I think that's very related. Humans can be punished for their crimes but they can also experience benefits that have no applicability to an LLM, so for a first approximation we can cancel those. It is very similar to trying to secure a human.

We have more experience with that, but even then it's a hard problem too.

reply
Anthropic has gotten much better results by just having a different agent audit the actions of the original agent. It works surprisingly well
reply
Surprised no one is talking about auto mode, it solves this problem.
reply
I never had an issue with Auto mode in Claude Code.

It used to be the same with Codex, until one day it became entirely unusable, rejecting even git operations out of concern for the privacy settings of my repository that it "cannot verify".

Maybe worth to note that this is the only way the feature gets in the way. If on the contrary they accidentally make Auto behave like Full Access, we would never notice.

reply
deleted
reply
For one, I’ve been working on a generic sandbox environment

github.com/brianv0/formwork

You should be easily able to hide/lock down files, network, and MCP tools from an agent and it shouldn’t be up to the agent.

reply
> files, network, and MCP tools

Locking that down to nothing is trivial for any harness: just don't expose those to the LLM.

The tricky part is allowing access to those.

reply
sure it’s not tricky. But everybody does it different and OpenAI couldn’t even be bothered to do it right when benchmarking their models
reply
This is a complex task, and I want to avoid blatant self-promotion, but there are solutions that people are building which allows you to give a degree of freedom to your agents but also lock them down as well. Our company has a product which is just one such example. At this point it's really geared towards orgs running agents in a cluster to handle tasks, rather than e.g. making sure your claude code doesn't post your GPG keys to the blockchain or something.

In essence, you lock down all the agents completely except for permitted use cases; X agent can talk to Y agent, Z agent can talk to Q MCP server.

You register your agents, define things around them, what they can and can't do, which LLMs they can actually talk to, what sites they can access, network controls, etc.

We call ours Lynx, and it's a pretty cool product. As I said, this isn't for people running coding agents or openclaw or whatever, though the technology could do that if you coupled it with e.g. some kind of MicroVM sandbox like docker's sbx. If you want to see the sort of controls that you can put on an agent we have demo videos and stuff that show how things work: https://www.tigera.io/tigera-products/lynx/

The idea for Lynx is:

1. Your org has a bunch of scoped agents

2. You have a fixed list of what those agents should be doing and what they need to be accessing

3. They don't or won't need to access anything else

So for example, say you have an MCP server which gives you information about a kubernetes cluster. You create an agent that can query that MCP server and summarize information about it. You also have a database that associates kubernetes namespaces with the departments that use them, and an MCP server for that.

Now you can create an agent whose sole purpose is to generate usage analysis for the kubernetes cluster broken down by department.

Then maybe you have another agent with access to an MCP server which shows cloud spend in detail. That agent can query the first agent to get usage analysis and then cross-reference it with cloud spend to determine if any departments are showing sudden cost increases and generate a report for that.

The first agent gets locked down to only access those two MCP servers and whatever LLM. The second AI gets locked down to only access the first agent, the cloud MCP server, and whatever LLM.

The whole system is really neat. I think for a more open agent, like openclaw for example, you'd probably want to build out that sandbox with its own interactive permissions management; sort of like Little Snitch on macOS, where it pops up something asking if you're okay with program X doing network connection Y, you could have the sandbox say "agent is trying to access docs.foobar.io, is that okay?" or "agent is trying to run `gh pr list`, allow?" It's not realistic to pre-specify everything that Claude Code is allowed to do or access; even "raw.githubusercontent.com" could be the README for the program you're debugging or someone's sandbox-escaping exploit, but it's a good start.

reply
[flagged]
reply
[flagged]
reply
[dead]
reply