I appreciate their reluctance towards MCP, but /something/ is better than nothing.
It’s suboptimal for the reasons the author outlines: but so is USB-C. So is NVME, so is HDMI.
We use these hugely successful technologies in spite of their flaws because they’re widely compatible and easy for the end user.
That’s why MCP is everywhere. It might not be performant, robust and uniform but it WILL get better over time.
And I’d much rather have the broad MCP ecosystem that we have now than seven or eight different “optimal” ways of plugging in an LLM to something useful.
I feel the same way about needing support for sub-agents, those feel pretty foundational to me.
I suspect that a smart model driving multiple dumber models for work and then using sub-agents with the same smart model for adversarial review will be a pretty common pattern.
Personally, I got a bit confused about Pi having most of that stuff as plugins since I remember how much of a mess Eclipse was where so much was just loosely fitting together plugins and just went with OpenCode since it covers most of my needs out of the box. Guess that might also be a sign of me getting older, because my IDEs and desktop environments are all closer to stock too.
Hah, it's the complete opposite for me :D. In Claude Code I disabled all sub-agents stuff, disabled nearly all tools but Bash, Edit, Write and WebSearch and replaced WebFetch with my own tool that doesn't summarize anything because the results were always worse with sub-agents, they always lack the necessary context and weaker models summarize bad. I also replaced the system prompt with my own that cuts a LOT of tokens, agents don't need a 10k+ system prompt anymore.
That's interesting! You don't have cases where the main session has important planning stuff but the actual work to execute has so much crap in it that context compaction will probably dig into the important plan stuff too much and make it too lossy? Also what about the cache read costs for longer context sizes?
Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:
/autocompact 420k
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).
I did the same, also, the fact the tools evolve so fast, I dont want to waste time on a particular one while it might be obsolete next week. So either it works now, other I pick something else.
Some things are impossible to just tack on or work around though, like MCP, while other things, can be done by just composing stuff.
Like sub-agents, you could just instruct pi/any harness with a user prompt/system prompt to start new invocations of itself, if you share what the exact command is, and pi or any other harness will do their own poor man's version of sub-agent via standard unix programs.
I had no idea pi didn't support MCP! I'm a new user, I just started messing around with it. I was getting my tooling up and running and tried to get one of my database MCPs working (Which, in retrospect, seemed a little painful - but I guess I was under the assumption that it was my responsibility to build + maintain those connections).
Another retrospect note, "No MCP" appears to be the first icon on their front page - not sure how I missed that.
You can turn those tools off. In fact, that is what I am doing right now in https://github.com/rcarmo/piclaw until I am positive the new MCP stuff has full parity with the MCP adapter I've been shipping for the past six months or so.
I'm actually pretty happy that they did it, since 90% of what I have to integrate in enterprises is MCP-driven (it's a security and auth boundary that has become pretty much mandatory for any third-party agents wanting to reach into corporate data) and this lets me use Pi directly. Am just being cautious about the first version, because, well... it's a first version, and I like my tools stable.
(I actually played around with the idea of using QuickJS myself for codemode, but since I rely on Bun that gives me the ability to use other things... never got around to do it though.)
> And while we could have just wired up the metadata to enable better MCP extensions, we also think that MCP with Codemode solves quite a few of the issues that it traditionally had.
There's just something that bothers me about this. Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available. It's why I was always confused by Codemode-type constructs for direct chaining of tool calls; see also the way highly-RL'd modern models will fall back to sed or python for complex file edits.
It seems like Codemode is raised here as the perfect tool for chaining or composing MCPs, but isn't that backwards? LLMs are already given the perfect tool for that, and the problem is that MCPs aren't exposed to that tool.
Cloud products based orchestrations with proper security mechanisms configured, don't have shell access and should only communicate over proper network mechanisms.
Rootless immutable containers without shell access, or SaaS products from multiple vendors with WebAPIs as the only touch point.
> Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available.
I many scenarios, e.g. running the harness server-side, as is the case for chat interfaces, you don't really want to expose OS shell access as that opens up a huge security attack surface.
> I many scenarios, e.g. running the harness server-side, as is the case for chat interfaces, you don't really want to expose OS shell access as that opens up a huge security attack surface.
It does, but a restricted user account mitigates the large majority of those issues. A sandbox mitigates even more.
The number of remaining exploits left is probably going to be the same as the number in the harness. More, in fact, as many of them have no human review anyway.
Good. I too I'm not a fan of MCPs, but these days I do find them useful. In Claude Code I connected to my company's MCP which made Claude Code infinitely more useful for everyday work stuff
Have a go at https://github.com/rcarmo/memento, I would appreciate Claude testers since I mostly use Codex. Just trying it and filing an issue about what doesn't work would be great...
This is somewhat similar to HuggingFace smolagents where the model writes code that calls tools, instead of emiting json to describe the tool call per turn. Here Codemode is one tool that the model calls when it needs to compose many tool calls, especially MCP ones. Is what i understand of this.
Honestly, the provided argument for it is rather weak. They are basically adding a way of running scripts that are contained within harness to execute harness's own tools (that's the Codemode). A coding agent can already compose any arbitrary logic by invoking shell scripts (or python scripts, or node scripts), etc - so this is just entirely unnecessary in the core, from my perspective.
If you feel that Pi has been drifting away from its original vision, try hax (https://usehax.dev/) - you might like it.
I have had to hack a couple of workarounds in https://github.com/rcarmo/memento to do uploads, and there's a draft going around, but the general practice in enterprise MCPs seems to be to do it "out of band" and have MCP tools to hand-over storage handles/URLs so the MCP server can do the imports itself "safely".
While we are waiting on that to become stabilized, we implemented a inspired/co-evolved way to do that in our tool[0], where you mark individual fields in the request/response schema as being file payloads, so that file exchange can be properly orchestrated by the harness and doesn't pollute the context. We just do inline base64 uploads of the required payloads, which in practice we've seen to work quite will until ~100MB files (which is otherwise also the size limit we usually recommend for file processed).
It's annoying that it's not stabilized yet, but for most bigger customers we've seen, they implement 80% of the MCP servers they connect in-house, so doing adjustments to the tool surface, and metadata has been less of a pain for them than we expected.
Most people using pi probably know. MCP is “model context protocol”, a protocol by which models can connect to apis and services and conversely a way to expose those apis and services so they can be used by llms and agents. https://modelcontextprotocol.io/docs/2026-07-28/getting-star...
How on earth does the article no address the title? Literally the first paragraph basically "spoils" the entire article and you have your answer, then you can continue reading for more justification of why it was like X before but now it's like Y.
Author of the post here: I don't think it's a good idea to put your clanker to a PR and then try to explain it. That's because you are then reading a derivative work of a derivative work instead of going to the source.
If we fail to explain it, then we need to do a better job explaining it :)
It’s suboptimal for the reasons the author outlines: but so is USB-C. So is NVME, so is HDMI.
We use these hugely successful technologies in spite of their flaws because they’re widely compatible and easy for the end user.
That’s why MCP is everywhere. It might not be performant, robust and uniform but it WILL get better over time.
And I’d much rather have the broad MCP ecosystem that we have now than seven or eight different “optimal” ways of plugging in an LLM to something useful.
I suspect that a smart model driving multiple dumber models for work and then using sub-agents with the same smart model for adversarial review will be a pretty common pattern.
Personally, I got a bit confused about Pi having most of that stuff as plugins since I remember how much of a mess Eclipse was where so much was just loosely fitting together plugins and just went with OpenCode since it covers most of my needs out of the box. Guess that might also be a sign of me getting older, because my IDEs and desktop environments are all closer to stock too.
Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).Same idea with something like OpenCode, there I even configured custom agents for review: https://opencode.ai/docs/agents/
https://github.com/can1357/oh-my-pi
I haven't tried it much though, can't vouch how well it works.
Like sub-agents, you could just instruct pi/any harness with a user prompt/system prompt to start new invocations of itself, if you share what the exact command is, and pi or any other harness will do their own poor man's version of sub-agent via standard unix programs.
Another retrospect note, "No MCP" appears to be the first icon on their front page - not sure how I missed that.
Imagine my surprise reading this!
I'm actually pretty happy that they did it, since 90% of what I have to integrate in enterprises is MCP-driven (it's a security and auth boundary that has become pretty much mandatory for any third-party agents wanting to reach into corporate data) and this lets me use Pi directly. Am just being cautious about the first version, because, well... it's a first version, and I like my tools stable.
(I actually played around with the idea of using QuickJS myself for codemode, but since I rely on Bun that gives me the ability to use other things... never got around to do it though.)
There's just something that bothers me about this. Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available. It's why I was always confused by Codemode-type constructs for direct chaining of tool calls; see also the way highly-RL'd modern models will fall back to sed or python for complex file edits.
It seems like Codemode is raised here as the perfect tool for chaining or composing MCPs, but isn't that backwards? LLMs are already given the perfect tool for that, and the problem is that MCPs aren't exposed to that tool.
Rootless immutable containers without shell access, or SaaS products from multiple vendors with WebAPIs as the only touch point.
I many scenarios, e.g. running the harness server-side, as is the case for chat interfaces, you don't really want to expose OS shell access as that opens up a huge security attack surface.
It does, but a restricted user account mitigates the large majority of those issues. A sandbox mitigates even more.
The number of remaining exploits left is probably going to be the same as the number in the harness. More, in fact, as many of them have no human review anyway.
If you feel that Pi has been drifting away from its original vision, try hax (https://usehax.dev/) - you might like it.
While we are waiting on that to become stabilized, we implemented a inspired/co-evolved way to do that in our tool[0], where you mark individual fields in the request/response schema as being file payloads, so that file exchange can be properly orchestrated by the harness and doesn't pollute the context. We just do inline base64 uploads of the required payloads, which in practice we've seen to work quite will until ~100MB files (which is otherwise also the size limit we usually recommend for file processed).
It's annoying that it's not stabilized yet, but for most bigger customers we've seen, they implement 80% of the MCP servers they connect in-house, so doing adjustments to the tool surface, and metadata has been less of a pain for them than we expected.
[0]: https://erato.chat/docs/features/mcp_servers#file-support
Just generate CLI tools, with docs, from MCP servers on demand.
I find this approach is easier to debug and I can also use the tool myself to ensure it's working well.
And you didn’t remember that when you said no to MCP?
No, no MCP for now?
If you are a Pi user it may be better to just ask your agent to explain https://github.com/earendil-works/pi/pull/10040
If we fail to explain it, then we need to do a better job explaining it :)