Together with my colleagues, I attended AI Engineer London this week. I had a blast: lots of new connections were made, and of course I enjoyed the buzzing city.

My personal takeaways:

Code mode over bloated MCP wrappers

Avoid exposing thousands of tools via an MCP server, which is itself mostly just a wrapper around your API. This is the current status quo for many companies. Instead, give it code mode, a term coined by Cloudflare, which allows your LLM to write small snippets of code that interact with your API and are executed in small isolates. This makes things more flexible, less LLM-heavy, and much faster, as it avoids going back and forth between the LLM and your API/MCP.

Local models are finally becoming usable

Local models are finally getting to the point where they are usable. With the recent release of Gemma 4 from Google DeepMind, this has been taken to another level. This model runs on smartphones, IoT devices, etc. For many tasks, we do not really need all the compute from frontier LLMs. Imagine asking what can be seen in a photo, enhancing your grocery list, or just having a smarter Siri on iPhones, please Apple. The second advantage, of course, is that it is all local.

MCP apps are coming fast

MCP apps are coming fast, and with this, maybe a glimpse into the near future of the internet. With this extension to the protocol, it is possible for vendors to ship their own look and feel of their apps into any MCP client supporting this feature. This is a win-win-win for the LLM providers, the software providers, and finally the user, as they do not have to leave the LLM interface and everything can possibly be done from one place. The smart thing about this integration is that the LLM stays informed about the user’s actions inside the MCP apps and therefore has the context to reason. Personally, I am really curious where this is going and whether it will be mass-adopted. To be honest, I have not yet heard of many people directly buying clothes or booking their holiday through, let’s say, ChatGPT.

Developer productivity: just do things, but do not rush

This point is more about developer productivity with LLMs. In short, you can “just do things.” Just do not fall into the trap of moving too fast. Systems architecture, choosing the right features to build, and the core architecture of your software, just to name a few, all need friction. The process of going back and forth and thinking about problems and solutions is needed. This might feel slow in the moment, but it is crucial for any successful outcome. Whether you create this friction by talking to an LLM, your coworker, friends, or by just taking a break instead of maxing out tokens does not matter as much, in my opinion. Slowing the fuck down, as nicely said by Mario Zechner, seems like a good thing in a time where many people just build things for the sake of burning tokens, building, and experiencing FOMO. Take a break in nature, take time to think outside the box, and enjoy life. Most of my best ideas, anyway, I had away from the screen.