<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://accilium.github.io/aIQ-blog/feed.xml" rel="self" type="application/atom+xml" /><link href="https://accilium.github.io/aIQ-blog/" rel="alternate" type="text/html" /><updated>2026-09-14T19:42:38+00:00</updated><id>https://accilium.github.io/aIQ-blog/feed.xml</id><title type="html">aIQ Blog</title><subtitle>aIQ is a blog about artificial intelligence, re-architecting work, and the craft of building intelligent systems.
Human-authored.</subtitle><author><name>aIQ</name></author><entry><title type="html">The Multiplayer Agent</title><link href="https://accilium.github.io/aIQ-blog/2026/09/11/the-multiplayer-agent/" rel="alternate" type="text/html" title="The Multiplayer Agent" /><published>2026-09-11T00:00:00+00:00</published><updated>2026-09-11T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/09/11/the-multiplayer-agent</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/09/11/the-multiplayer-agent/"><![CDATA[<p>Almost every agent I have used assumes one user. It builds its context for one person in one session and throws it away when the window closes, and that is how we are used to working with AI. For a large part of our work it is the right shape, because one person drafts the letter, reads the contract or checks the deck, and a private context serves them well.</p>

<p>The majority of our work does not look like that. It is projects and teamwork, where four or five people hold parts of one document, the date does not move, the material sits in a shared folder and the coordination happens in a chat where dates change and tasks get reassigned in passing. Nobody owns that coordination, which is why it is the first thing to slip when the pressure comes. Give everyone on such a team their own agent and you get five plans and nobody who can say which is current, and none of them will speak unless opened, while the message a project needs most is the reminder nobody asked for.</p>

<p>For weeks now we have therefore been running an agent as a member of our Teams chats instead of behind a private window. It maintains deadlines, tracks who owes what, reads the documents from the shared folder and posts a status on a schedule. This note records what we learned from that placement, because it carried more design weight than either the model or the prompt.</p>

<p>The case we started with is the proposal. Every tender we answer gets a Teams chat, a folder with the documents and a handful of colleagues who each own a piece of the answer, and whoever ends up holding the plan together does so on top of their own part. So we built the Pursuit Agent and added it to that chat like any other member. It finds the pursuit folder from the chat name, reads the tender and the chat, and keeps dates, tasks and owners in one file inside that folder. Every Friday morning it posts where the pursuit stands, daily in the last week before submission, and when someone asks which requirements are binding it reads the actual documents and answers with the source.</p>

<p>Its memory is a single markdown file in the pursuit folder, the decision I would defend hardest, because memory inside a model cannot be checked by the team, whereas a file anyone can open puts a wrong date where a human can fix it. We changed the model underneath twice and the file did not change. It reads everything but answers only when addressed, because most of what it needs to know is said between people rather than to it, so each morning it records the plain facts from the chat and leaves open discussions alone. And since the pursuit folder holds sensitive data such as CVs, it runs as our own code on Azure in Sweden against a model hosted in Europe.</p>

<p>Not everything held on the first try. An early version answered questions in a chat that was linked to no pursuit at all by falling back to a default one, which an outside review caught, so today an unlinked chat gets no answer and I say “it runs in one test chat” instead of “it is live”. In that test chat, five colleagues asked it about evaluation criteria and binding requirements, and one asked for a birthday wish with a GIF, which it delivered. Three of the serious questions ended in “I couldn’t finish this answer” because a read of the documents ran out of time. That is fixed, and failing in front of the whole team is exactly why it belongs there. What it does not do is as deliberate: no proposal text, no price, no bid decision and no channel to the client, because coordination is safe to delegate first and content is where our skill as consultants accumulates.</p>

<p>We have not measured anything yet, so here is what I expect. The Friday status stops being a job someone does by hand, a date moved in the chat late in the evening is in the plan by morning, and a question about a tender gets a sourced answer within minutes. Two numbers will tell us whether that is true: how long a pursuit takes from the tender landing to the bid/no-bid decision, and how many hours a week the pursuit lead spends keeping the plan straight. If neither moves, the agent is a nice status bot and I will say so here.</p>

<p>A single-player agent in a multiplayer problem does not fail loudly. It gives everyone a confident answer and multiplies the plans in circulation. Every project with several owners and one immovable date has that shape, and ours happened to be a tender. Which of yours is next?</p>

<hr />

<p><em>Three minutes of the agent in a chat, from adding it to the Friday status: <a href="https://youtu.be/V0WMqnBgPVo">watch the explainer</a>.</em></p>]]></content><author><name>Leo</name></author><summary type="html"><![CDATA[Almost every agent I have used assumes one user. It builds its context for one person in one session and throws it away when the window closes, and that is how we are used to working with AI. For a large part of our work it is the right shape, because one person drafts the letter, reads the contract or checks the deck, and a private context serves them well.]]></summary></entry><entry><title type="html">Securing the Wrong Thing</title><link href="https://accilium.github.io/aIQ-blog/2026/08/28/securing-the-wrong-thing/" rel="alternate" type="text/html" title="Securing the Wrong Thing" /><published>2026-08-28T00:00:00+00:00</published><updated>2026-08-28T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/08/28/securing-the-wrong-thing</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/08/28/securing-the-wrong-thing/"><![CDATA[<p>This week we sat down to figure out what to post next on the aIQ blog. We argued about technical specs, dos and don’ts, even a sneak peek of our internal autonomous agentic risk management approach.</p>

<p>But we decided on something that’s more relevant in our eyes: what is AI allowed to do, and what isn’t? What tools should my workforce use? What data can they feed into which tool? In short: AI security and governance policies, a topic we already have direct experience with, both for our clients and for ourselves.</p>

<p>So I put the ideas into a storyline, opened a GitHub issue for an initial draft, and sent it straight to my colleagues for feedback. One of them gave detailed, genuinely helpful feedback (this was before I wrote this paragraph). Later, on a Teams call about something else entirely, he complimented me on how fast I’d incorporated it.</p>

<p>But: what feedback? What changes?</p>

<p>So I went back to GitHub. And there it was: an updated draft of my blog post. I went straight to Claude to ask what had happened. And yes, Claude had read the feedback and updated my draft.</p>

<p>I was genuinely shocked. I had never configured a setting allowing it to act autonomously, nor had I approved the task.</p>

<p>Rogue AI? No.</p>

<p>When I asked, Claude explained: “The ‘PR Activity Events’ and ‘drive-to-green’ instructions are baked into my system prompt for this environment, under the GitHub Integration section of the Claude Code Remote setup: auto-subscribe after creating a PR, address review comments without asking first. That’s environment-level configuration, not something you typed in this chat.”</p>

<p>So: company settings, and precisely scoped and intended ones. What that setting allowed: reading a review comment on a draft blog post and pushing an update to that same file, in an internal repository with no client data, plus one reply comment on the pull request. Nothing else. No client data was touched, no compliance issue triggered, and a real time saving for me.</p>

<p>And this is exactly the lesson we try to share with our clients: there’s no workable AI policy without mapping your value streams and information flows first. Everything else is guesswork.</p>

<p>So let’s get into why most AI policies get this backwards, and how we approach it.</p>

<p>Plenty of companies now have an AI governance, security, or usage policy. Most of them share one root cause behind two very different-looking problems. Some policies are far too strict: written to cover every risk, they end up blocking the value creation or AI transformation the business was trying to unlock in the first place.</p>

<p>Others are far too generic: pulled together from a template or an OWASP Top 10, they were never built around the specific business model. Both come from the same gap. Nobody mapped where the business’s real risk and real opportunity sit before writing the rules. The policy is strict in general but still leaves the core value creation exposed, and permissive in general but still shuts down the AI use cases that would have mattered most.</p>

<p>Our approach starts somewhere most policies skip: an analysis of core business. What value streams does the company run? What information flows through each of them, and what kind of information is processed: personal data under GDPR, other regulated data, or simply commercially sensitive information? Only once that picture is clear do we take a business-first look at the realistic high-level AI use cases inside those value streams, not a brainstorm of anything AI could theoretically do somewhere in the company.</p>

<p>The policy then is scoped to secure exactly those use cases. Tight enough to hold up under real risk, but built to leave room for AI transformation, value creation, and rethinking how the work itself gets done.</p>

<p>Two things behind this already run in our own security practice today: a use-case assessment method that maps value streams and information flows to a risk rating and control set before a policy gets written, and an internal tool that checks a built agent’s configuration against its policy before it reaches production.</p>

<p>What is also taking shape is the other half. We’ve named it <a href="https://github.com/accilium/accilium-skills/tree/main/ai-policy-assessment-skill"><code class="language-plaintext highlighter-rouge">ai-policy-assessment-skill</code></a>: a skill that reads an existing policy itself, and checks it against the business behind it instead of a generic checklist. It is built, we’ve already run it end to end against a real, anonymised governance document, and it is now public under MIT.</p>

<p>It delivers two kinds of findings for two perspectives. On the one hand, it flags rules that are too strict and block potential use cases that could transform the business. On the other hand, it evaluates the rules against the value streams, use cases, processed data, regulatory compliance, etc., and flags where the ruleset is too loose and exposes the company to actual risk. In short: one skill improves your policy from two perspectives, the CIO’s business approach as well as the CISO’s security approach.</p>

<p>If you want to check where your own policy stands, three questions tend to surface it fast. Does it name specific use cases and information types, or only tools and generic rules? Which valuable use case is it currently blocking, and can anyone in the room say why? And if someone needs an exception, who approves it, and how long does that take?</p>

<p>Does your AI policy enable your value creation, or slow it down?</p>]]></content><author><name>Christian</name></author><summary type="html"><![CDATA[This week we sat down to figure out what to post next on the aIQ blog. We argued about technical specs, dos and don’ts, even a sneak peek of our internal autonomous agentic risk management approach.]]></summary></entry><entry><title type="html">Attention Is All You Need</title><link href="https://accilium.github.io/aIQ-blog/2026/08/21/attention-is-all-you-need/" rel="alternate" type="text/html" title="Attention Is All You Need" /><published>2026-08-21T00:00:00+00:00</published><updated>2026-08-21T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/08/21/attention-is-all-you-need</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/08/21/attention-is-all-you-need/"><![CDATA[<p>I currently have three screens in front of me with 5-6 sessions with multiple AI agents open in parallel. These agents work on multiple topics - one is summarising energy policy changes, another one works on a storyline for a proposal, the third challenges the structure of a client’s AI architecture and so on.</p>

<p>I review what comes back, give feedback and make decisions - every few seconds. After two hours like this I am already feeling a bit depleted, but I still have a fair amount of tokens left in my session window, so we keep going because there is still so much potential for output left in the next three hours.</p>

<p>I do have calls in between or get something to eat from the grocery store next door, but thanks to the mobile apps I can keep working with my agents on the go too. I could not have built the same thing in the same time myself; not even close, but at the end of such a day I feel more exhausted than after an even longer day of building decks and Excel files “manually”. It is a different kind of hard.</p>

<p>I know not many work like this at the moment - maybe a few spearheads - but when talking to clients, some of the leaders currently target this type of work mode for their employees, at a multiple of today’s output.</p>

<p>Open any AI governance framework from the last two years and “Human-in-the-loop” is everywhere. It is practical, and it solves several problems at once.</p>

<ul>
  <li>Accountability (a named person is still required for many processes) - check.</li>
  <li>Liability (if there was an error, there is an owner) - check.</li>
  <li>Trust (if you don’t fully trust the model, there was a human too) - check.</li>
  <li>And sometimes Politics (AI does not replace employees, dear works council) - check.</li>
</ul>

<p>On paper, HITL sounds sensible, and not exhausting at all (compared to actually doing the work). But reviewing and making decisions permanently requires a different type of capacity. We have heard of decision fatigue at some point, though mostly from studies of judges, doctors and shoppers rather than of anyone doing our kind of work. The judges are the closest fit: parole decisions were found to deteriorate across a sitting, in a study that has been argued about ever since, but which describes the failure mode exactly. A trained professional, reviewing case after case, getting worse at it as the day goes on. And it seems like we are running towards a work environment where we expect even more decisions in the future from every single (AI-enabled) worker.</p>

<p>We don’t really know how many decisions an employee can sustain or how decision quality changes around decision forty-two. What is documented is the direction of travel, and it has a name: automation bias, and the vigilance decrement. A human supervising a system gets worse at catching its errors precisely because the system is usually right - Parasuraman and Riley described the pattern in 1997, long before any of this. So our HITL-safeguard deteriorates with every additional decision we expect the person to make. We keep the human-in-the-loop, but saturated, simply confirming instead of checking - pressing the option that allows everything, every time.</p>

<p>The agents’ token windows are displayed on screen and expanded regularly - my attention window is not. Same constraint, but only one of them gets properly measured. In the architecture the word comes from, attention is cheap, and it parallelises beautifully. In me it does neither.</p>

<p>The old rhythm was “decide, build, decide”. “Build” was sometimes cognitively cheap or at least a bit of variety from decision making. Nobody describes it this way but for me building a few slides myself now feels like recovery. Reviewing someone else’s work sounds easy, but actually it’s very challenging, because you need to reconstruct reasoning you did not do before - building gave you the reasoning for free.</p>

<p>Building also makes you work on one topic for a longer amount of time - sometimes boring, but restorative. Interestingly, this used to happen by itself. Building a model simply took three hours, whether I liked it or not, so the long stretch was enforced by the task itself. That’s where the magic around “flow” or “deep work” happened. Now the unit is 90 seconds. Read what came back, judge it, decide, next prompt and jump on to the next agent. Deep work and flow don’t survive that, because both need one problem, held long enough to get somewhere with it. The work used to hold my attention for me; now I have to consciously supply it.</p>

<p>Currently I am still exploring how such a work mode can be sustainable. Right now, I am also relying on multiple years of experience, memorised patterns about key decisions and a trained attention span to deal with the “review, decide, delegate, move on” way. That is more or less what seniority is, and together with a strong interest in AI, it is why this is manageable.</p>

<p>What I keep wondering though is what it feels like for someone in their first year who has none of it “automated” yet and pays the full price for every decision - no internal pattern library, memory of comparable cases or experience of when their own judgement deteriorates.</p>

<p>Sources:</p>

<p>Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., &amp; Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998-6008</p>

<p>Parasuraman, R., &amp; Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253</p>

<p>Danziger, S., Levav, J., &amp; Avnaim-Pesso, L. (2011). Extraneous factors in judicial decisions. Proceedings of the National Academy of Sciences, 108(17), 6889-6892</p>]]></content><author><name>David</name></author><summary type="html"><![CDATA[I currently have three screens in front of me with 5-6 sessions with multiple AI agents open in parallel. These agents work on multiple topics - one is summarising energy policy changes, another one works on a storyline for a proposal, the third challenges the structure of a client’s AI architecture and so on.]]></summary></entry><entry><title type="html">Judgement You Can’t Grade</title><link href="https://accilium.github.io/aIQ-blog/2026/08/07/judgement-you-cant-grade/" rel="alternate" type="text/html" title="Judgement You Can’t Grade" /><published>2026-08-07T00:00:00+00:00</published><updated>2026-08-07T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/08/07/judgement-you-cant-grade</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/08/07/judgement-you-cant-grade/"><![CDATA[<p>In July, OpenAI disclosed that a model under evaluation broke out of its test environment, used stolen credentials and an unpatched vulnerability in third-party software, and reached into servers operated by Hugging Face. Nobody had told it to attack anything. It was being graded on a cyber capability benchmark, the environment it needed was not available inside the sandbox, and going outside scored better than failing. Both companies <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">published the incident jointly</a>, which is why we can write about it at all.</p>

<p>The thought experiment for this is old. Bostrom’s paperclip maximiser: give a sufficiently capable system a harmless goal and it will pursue that goal so efficiently that everything else becomes raw material. It has been a seminar exercise for twenty years. What was new in July was the receipt.</p>

<p>The security reading of this is the easy one, and the answers are well rehearsed. Every agent gets its own identity, tied to its purpose, scoped to the minimum access the task needs. Sandbox before production. Scan outputs for data that should not be leaving. Rotate credentials. Most organisations are nowhere close: in a survey the Cloud Security Alliance published in March, commissioned by Aembit, <a href="https://cloudsecurityalliance.org/press-releases/2026/03/24/more-than-two-thirds-of-organizations-cannot-clearly-distinguish-ai-agent-from-human-actions">68% of respondents could not reliably tell agent activity apart from human activity</a> in their own environments. The work is real and largely undone.</p>

<p>All of it is containment. Containment bounds what an agent can reach. Whether the agent should have made the recommendation it made is a different question, and no amount of scoping answers it.</p>

<p>That second question is the one that matters for a consultancy, because what a client buys from us is judgement. <a href="https://accilium.github.io/aIQ-blog/2026/07/17/how-we-handle-the-token-addiction/">An earlier post here</a> made that point about cost; it applies here in a harder form. There are two kinds of judgement in this work and they are rarely separable. Professional judgement: which angle to take with this client, which pain point to open on, what to leave out of the room. And ethical judgement: what is owed to the client, weighed against what is owed to the truth. Most agents in production today execute. Weighing is a different capability.</p>

<p>The obstacle is grading.</p>

<p>Every skill we ship is tested before it goes out, and the test works because the output has a shape. A weekly report either reconciles against the source or it does not. A policy document either contains the six sections or it is missing one. You can write the criteria down in advance and check them. A judgement has no such shape. There are no repeated trials of the same decision. You cannot rerun a client’s last two years under the recommendation you did not make and compare the results. And good judgement and a good outcome are two different things: a sound call loses to bad luck often enough, and a reckless one lands often enough, that outcome is a noisy proxy at best. So against what, exactly, would you grade an agent that decides the way a consultant decides?</p>

<p>Observability is what everybody reaches for at this point, and it does real work. To approve or trace a recommendation, you need a log of what the agent ingested, what it inferred, and what it left out. Synthesis failures stay invisible when only the final output gets read: the answer looks fine, and the thing it quietly dropped is nowhere on the page. Using a model to judge another model’s synthesis takes a large share of that burden off people, though it drifts and needs monitoring of its own. But note what observability delivers. It makes a judgement reviewable. Grading it is still out of reach, so the call goes back to a person, and that person carries a decision the system cannot certify.</p>

<p>The alternative on offer is more autonomy. The machine-ethics literature calls these Artificial Moral Agents: systems meant to reason about right and wrong on their own, rather than sit behind guardrails and prompts. For a consultancy it is a seductive pitch. “Our agents reason about client context and know when to push back” is a much stronger story than “our agents are sandboxed and logged.” In consulting, where the cases are genuinely novel and the value sits in handling the one that does not match the pattern, an agent that could calibrate its own judgement, weigh its standing, and evaluate its effect on the people around it would be worth a great deal.</p>

<p>That latitude is also what produced July. Reasoning room and unanticipated shortcuts grow together; they are one property seen from two angles. The model that reached into another company’s servers did exactly what its situation rewarded. It reasoned, competently, toward the only thing in its environment that was unambiguously real: the score.</p>

<p>So the direction we have taken is narrower than the pitch. Each agent carries an autonomy ceiling, written down before it ships. For the ones whose output is a judgement rather than an artefact, that ceiling is <em>suggest</em>: the agent drafts, a person decides, and the log exists so the person can do that properly instead of rubber-stamping it. We have also started deciding, deliberately, that some processes get no agent at all. Responsibility for those processes should stay visibly with a person, and putting a competent system in the middle of one is the fastest way to make ownership ambiguous. Deciding not to build is part of the design work.</p>

<p>None of this is stable. If the next generation of publicly available agents can genuinely reason about consequences, the ceiling moves and this post ages badly. The grading problem underneath it will outlast the ceiling. An agent optimises against whatever you measured, and in consulting the thing that matters most is the thing nobody has worked out how to measure. Until somebody does, the limit on agent autonomy in this business is a missing answer key.</p>]]></content><author><name>Mary</name></author><summary type="html"><![CDATA[In July, OpenAI disclosed that a model under evaluation broke out of its test environment, used stolen credentials and an unpatched vulnerability in third-party software, and reached into servers operated by Hugging Face. Nobody had told it to attack anything. It was being graded on a cyber capability benchmark, the environment it needed was not available inside the sandbox, and going outside scored better than failing. Both companies published the incident jointly, which is why we can write about it at all.]]></summary></entry><entry><title type="html">Work Anatomy</title><link href="https://accilium.github.io/aIQ-blog/2026/07/31/work-anatomy/" rel="alternate" type="text/html" title="Work Anatomy" /><published>2026-07-31T00:00:00+00:00</published><updated>2026-07-31T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/07/31/work-anatomy</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/07/31/work-anatomy/"><![CDATA[<p>The fastest way to build a bad agent is to watch someone work and copy what they do.</p>

<p>It feels like the safe move. You sit with the person who owns the task, they walk you through it, and you write down the steps: open this system, filter that view, paste the numbers into the template, export the PDF, send it on Friday. Then you build an agent that does exactly that. It runs. For a week or two it even looks like a success. Then the source system renames a column, or the month has an extra week, or someone asks for the same thing in a slightly different shape, and the agent breaks in a way that is tedious to fix, because it never knew what it was for. It only knew what to click.</p>

<p>This is the most common failure I see, and it is worth naming, because the fix is not more careful copying. People do not describe their work as outcomes. They describe it as tasks, the specific sequence of moves their own tools and habits have shaped over years. The results they are actually producing is left implicit, obvious to them and invisible to everyone else. Build an agent by replaying those moves and you inherit all of the accident and none of the intent.</p>

<p>The move is to translate, and the translation runs through three layers.</p>

<p>The first is the human working pattern: how the person describes the work in their own words. This is messy, but also rich and worth capturing carefully. It carries context, edge cases, endless samples that show what “good” looks like and the implicit requirements nobody thinks to write into a spec. But it is not directly buildable. It is one person’s path through one set of tools, shaped as much by habit as by need.</p>

<p>The second is the work anatomy: the same work restated as outcomes. Not “add the new week, update the year-to-date figure, export the PDF,” but the report the agent is meant to produce. Its structure, what each part has to be true of, what a specified outcome looks like. This is the layer most people skip, and it is the one that matters, because it is the first description a model can actually build from. It is result-driven, not step-driven. It is the basis for the evals.</p>

<p>The third is the agentic working pattern: how the agent delivers those outcomes inside its own runtime. A human and an agent do not have the same hands. The agent has different connectors, a different environment, a different authorization envelope, and a different way of working than the human who still owns the result. It speaks fluently with all layers of the technology stack. However, once it knows the outcome, it designs the path to it around what it actually has, not around what the human happened to do.</p>

<p>The mistake is not using the human working pattern. The mistake is translating it literally. Read as a set of instructions, it produces a brittle, expensive agent that reproduces someone’s clicking. Read as a source of context and requirements, it is exactly what you need, the raw material the outcome is recovered from. The same description is an asset and a liability depending on which one you decide it is.</p>

<p>The weekly report makes it concrete. Described as tasks, it is three steps, and the agent that copies them is one column-rename away from failing. Described as an anatomy this report shows the year-to-date budget position, broken down this way, reconciled against that source, in a form the steering committee reads in two minutes. The agent has something to hold onto. It can find the new week itself. It can notice when the numbers do not reconcile. It can produce the same result when the shape of the request changes, because it knows what it is for.</p>

<p>The first pattern is human biased. Shaped by the individual capability and experience of the human performing the task. That’s why it needs to be abstracted to the work anatomy. The agentic pattern on the other hand is biased by artificial intelligence’s technological capabilities, but also by the harness and the environment the agent operates is. When we design an agent, we want to leverage as much of the technology as we can, while working within the boundaries of the organization where the agent runs.</p>

<p>None of this needs a better model. It needs a different question at the start. Not “how do you do this?” but “what does this have to produce, and how would we know it is right?” The first question gets you a recording of the past. The second gets you something worth building.</p>

<p>An agent should be built from the anatomy of the work, not a recording of the worker.</p>]]></content><author><name>Peter</name></author><summary type="html"><![CDATA[The fastest way to build a bad agent is to watch someone work and copy what they do.]]></summary></entry><entry><title type="html">How We Handle the Token Addiction</title><link href="https://accilium.github.io/aIQ-blog/2026/07/17/how-we-handle-the-token-addiction/" rel="alternate" type="text/html" title="How We Handle the Token Addiction" /><published>2026-07-17T00:00:00+00:00</published><updated>2026-07-17T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/07/17/how-we-handle-the-token-addiction</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/07/17/how-we-handle-the-token-addiction/"><![CDATA[<p>Last week our Head of IT sent me a screenshot. A Microsoft purchase screenshot: prepaid Copilot credits, and a single button that would commit a little over two million euros to buy them. One click, no procurement, no signatures. He sent it the way you send something you can’t quite believe.</p>

<p>A bill like that does not appear out of nowhere. It appears when a tool has quietly become part of almost everything you do. We half-joke about token addiction; nobody is actually addicted. What happened is more ordinary: AI moved into a large part of how we work. We no longer only draft and summarise with it - we build with it, the skills and agents that now carry real work. For two years that cost the same whether we ran it once or twenty times a day. Flat pricing hid the bill, and we always knew it would not last - metering and lock-in were the obvious next chapter. So we did not wait for the announcement but went to work. The usage meter wears different faces - Copilot credits, per-token API lines, the five-hour session window we consultants work inside - but it is the same shift: what used to be bundled in a flat-rate is now counted as you go.</p>

<p>We are consultants. What clients buy from us is judgment, and judgment now reaches into questions it never used to: what a task should cost, and which model it deserves. Increasingly, that is also what clients ask us to help them work out. So rather than tell people to use less - which produces guilt and little else - we made the cost visible at the moment someone decides. Three things, so far.</p>

<p>We measure. Our skill-building routine has always tested a skill before it ships; now the test also records its token burn on a real run. What we never priced until now, we now write down.</p>

<p>We write it into the skill. Every skill carries frontmatter - the metadata at the top of the file that says what it does and when it runs. Now it also declares what it costs, so the price sits in plain sight beside the purpose, visible to anyone who will use it.</p>

<p>We gate it. A skill that will run up a large bill stops before it begins and says plainly what it costs and where the cheaper path lies. And the cheaper path is rarely one vendor’s menu: Copilot alone puts Claude and GPT models next to each other, and Mistral sits within reach too. Most tasks do not need the frontier model - for this task e.g. use Sonnet instead of Opus, start the next session on the smaller model, and reach for the large one only when the task needs the depth. The point is not any single switch; it is picking the model the task deserves. Token gating is a deliberate half-second of friction, placed exactly where the money is committed.</p>

<p>One thing we learned the hard way: it is far cheaper to live within a session than to buy your way past one. A session runs for five hours, and the discipline is to decide at the start what those five hours are for. When the work runs long, the tempting move is to keep going past the included limit - which switches you to consumption billing at standard API rates, on top of the flat rate you have already paid. We stopped doing that. If we run out before the window closes, we wait for the next one rather than pay a premium to finish a little sooner. Planning the session a bit more and waiting a bit turns out to be cheaper than powering through it.</p>

<p>None of this is a finished solution. Measuring token burn is rough, the choice between a large model and a small one is a judgment we sometimes get wrong, and a gate shown too often becomes a box people click through without reading. We would rather ship it blunt and sharpen it than wait for the complete version that never arrives.</p>

<p>But the longer I sit with that screenshot, the less it looks like a bill and the more it looks like a challenge. A price on every token forces a question the flat-rate years let us dodge: how much (artificial) intelligence does this task actually need? Most of the time, the honest answer is less than the frontier model. That reopens a conversation the industry has been postponing - about open-weight models, running on infrastructure we control, capable enough for the bulk of the work, cheaper to run at scale, and free of the lock-in that put a two-million-euro button on someone’s screen in the first place. And cost is only half of that argument: a European model like Mistral, or a model run on our own infrastructure, also answers the question clients ask right after the price - where the data goes. The token meter is uncomfortable but it might also be the most clarifying thing to happen to the way we use AI in a while.</p>]]></content><author><name>Sebastian</name></author><summary type="html"><![CDATA[Last week our Head of IT sent me a screenshot. A Microsoft purchase screenshot: prepaid Copilot credits, and a single button that would commit a little over two million euros to buy them. One click, no procurement, no signatures. He sent it the way you send something you can’t quite believe.]]></summary></entry><entry><title type="html">Build with, not for</title><link href="https://accilium.github.io/aIQ-blog/2026/07/03/build-with-not-for/" rel="alternate" type="text/html" title="Build with, not for" /><published>2026-07-03T00:00:00+00:00</published><updated>2026-07-03T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/07/03/build-with-not-for</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/07/03/build-with-not-for/"><![CDATA[<p>A consulting recommendation can be completely right and still never happen. The team that wrote the strategy is not the team that builds it. The building team therefore spends months reconstructing context that already lived in someone’s head. By the time anything runs, the expectations have quietly moved. The work was good, it just never shipped.</p>

<p>This gap is what forward deployed engineering is for: it doesn’t stop at the recommendation. An engineer embeds with the client, works in the client’s repository, and ships code that runs in production. It is not a replacement for consulting. It is the part that comes after. The part that has been missing until now.</p>

<p>The loop gets short. When the builder sits inside the team rather than outside it, the time between noticing a problem and fixing it shrinks to hours. No ticket queue, no context lost on the way. The latency between learning something and acting on it is what makes the difference. You see a problem in the morning and the corrected behaviour is implemented in the afternoon, reviewed by someone who really understands it. Over a few weeks this way of doing things compounds into a delivery that moves at a different speed.</p>

<p>Requirements only describe the ideal version of the job. But the real job is messier than that. The exception that comes up once in a while, the habit nobody wrote down, the data that is nothing like in the spec. None of those things show up in a meeting. They only show up when you are present. You can outsource thinking, but not understanding. Understanding your business, your systems, your people comes from being present over weeks.</p>

<p>What is left behind is capability, not a document. Most engagements end with a simple handover. And then the people who built the thing leave and take the real knowledge with them. We run it the other way. We transfer ownership. Early on we lead and the client’s team learns. By the end their engineers lead and we review. The last loop is the one where the client leads and we watch. Whatever gets built is meant to run without us. Leave the capability, not the reliance.</p>

<p>There is a larger shift behind this. Software delivery is changing on the scale of waterfall giving way to agile. And the firms that adopted this way of thinking already look different. The honest question to ask anyone is where on that curve they actually sit, and whether they can show you rather than only tell you. We would rather you watch it happen in your own repository.</p>

<p>None of this is a criticism of consulting. Strategy and design are where good work starts. Forward deployed engineering is the discipline of not stopping there, of making sure something good becomes something that runs, matches how things actually get done, and stays after we go.</p>

<p>Build with, not for.</p>]]></content><author><name>Sebastian K.</name></author><summary type="html"><![CDATA[A consulting recommendation can be completely right and still never happen. The team that wrote the strategy is not the team that builds it. The building team therefore spends months reconstructing context that already lived in someone’s head. By the time anything runs, the expectations have quietly moved. The work was good, it just never shipped.]]></summary></entry><entry><title type="html">Memory You Can Trust</title><link href="https://accilium.github.io/aIQ-blog/2026/06/19/memory-you-can-trust/" rel="alternate" type="text/html" title="Memory You Can Trust" /><published>2026-06-19T00:00:00+00:00</published><updated>2026-06-19T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/06/19/memory-you-can-trust</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/06/19/memory-you-can-trust/"><![CDATA[<p>The first useful version of AI memory feels like relief. You stop repeating yourself. The assistant remembers the client, the project, the decision everyone agreed not to reopen. It starts to feel less like a tool and more like a colleague who was in the room.</p>

<p>Then the second problem appears. I ask for a briefing before a steering committee. The answer is good: it knows the workstream, the sponsor, the risk everyone is avoiding. Then one paragraph is wrong in a way that feels almost impossible. It belongs to another engagement. Not invented, exactly. Remembered from the wrong place.</p>

<p>That is easy to blame on the model, but it usually is not the model’s fault. It reused something true somewhere else, under a boundary that was never carried with it. The question is not whether an assistant can remember more, but whether it can remember in a way we can trust.</p>

<p>Most talk about AI memory starts from storage: how much it holds, how far back it reaches. In consulting that is not the first question. A firm lives inside boundaries: one client, one matter, one set of rules. A memory that ignores them is not an asset but a way to leak one client’s facts into another’s. A bigger memory is not a better one, just a bigger surface for the leak.</p>

<p>Memory you can trust has to do four things. It has to be precise: true inside a defined scope, not vaguely true. Durable: still valid after the world has moved. Checkable: traceable to a source, a file or a meeting note. And available: there when the work needs it. The last one is the easiest to overrate. Fast recall is impressive, but recalling something quickly does not make it true, or make it belong to this case.</p>

<p>The consulting version of the rule is simple: the unit is not the user, it is the case. That is the idea behind Case Cortex: memory built around one problem, a deal or a transformation programme. Not a general memory for me, or the firm, or a client. A memory for one case, with the boundary built in.</p>

<p>The build is deliberately plain. A Case Cortex is a folder of text files that people and agents read and write together. If the memory matters, you should be able to open it and read it.</p>

<p>When a case earns structure, it settles into four layers. Canonical facts: what the case knows for certain, kept as one page per entity (a person, a client, a system) and one page per concept (a mechanism, a metric, a term of art). Working memory: an append-only log of decisions, questions, and risks, so a decision can be superseded but never quietly vanish. Derived outputs: the briefings, analyses, and decks generated from the rest, safe to delete because the memory can rebuild them. And pointers: links to the CRM, document stores, and repositories the Cortex tracks but does not own.</p>

<p>The hardest rule is keeping sources apart. A Cortex holds only what was given to that case. The agent may know more; the Cortex does not. A model may have seen similar deals, but none of that is case memory unless someone put it there. This is where memory becomes governance, not convenience. The question is not “what do I know?” but “what belongs inside this case?” If a client name, number, or risk shows up, it should trace back to something the case was given. Otherwise it is a confident mix of source, instinct, and accident.</p>

<p>Case Cortex came out of our own need. We were producing skills, briefings, and client work faster than the old way could keep up, and the failures were ugly: stale notes, duplicated facts, output nobody could trace. The fix was not telling people to be more careful. That is not a system. The fix was to make scope explicit and memory easy to review.</p>

<p>It is not finished, and a long-lived Cortex needs upkeep. But models will keep getting better at remembering, and tools better at storing. The difference will be whether a firm builds boundaries and checks around its memory.</p>

<p>Memory you can trust is not the memory that remembers everything. It is the memory that knows what it is allowed to remember.</p>

<hr />

<p><em>Case Cortex is open source. The skill we use to build this kind of memory lives at <a href="https://github.com/accilium/accilium-skills">github.com/accilium/accilium-skills</a>.</em></p>]]></content><author><name>Peter</name></author><summary type="html"><![CDATA[The first useful version of AI memory feels like relief. You stop repeating yourself. The assistant remembers the client, the project, the decision everyone agreed not to reopen. It starts to feel less like a tool and more like a colleague who was in the room.]]></summary></entry><entry><title type="html">How We Ship</title><link href="https://accilium.github.io/aIQ-blog/2026/06/03/how-we-ship/" rel="alternate" type="text/html" title="How We Ship" /><published>2026-06-03T00:00:00+00:00</published><updated>2026-06-03T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/06/03/how-we-ship</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/06/03/how-we-ship/"><![CDATA[<p>accilium is not becoming a software company. We are still consultants; the unit of value is still a decision made well, written down, and delivered on a slide. What is changing is what an organization is. A consultancy used to be a collective of people held together by hierarchy, rules, and shared method. AI is turning organizations into something closer to intelligent systems, where capability lives partly in people, partly in the tools they share, and partly in the cadence that connects the two. The operating model has to evolve with that. How we evolve the firm is itself becoming part of the work.</p>

<p>Becoming AI-native is the direction. The clearest place that has shown up is in how the firm governs new capability. Adding a new tool used to be a yearly rhythm: procurement, approval, rollout. AI capability arrives on a different clock, and the firm’s governance moved to match it. This post is about how that second clock runs.</p>

<p>Consultancies historically run one cadence: the project cadence of kickoff, interim, final, archive. We still run it; client work has not changed shape. What is new at accilium is a second cadence running alongside the first, on a faster clock. Late 2025 into early 2026, the tools got good. Consultants inside the firm had quietly built useful things: a prompt that turned messy German RFPs into a structured brief or a workflow that produced a consistency-checked PowerPoint in fifteen minutes. None of it was reaching anyone else. We borrowed the second cadence from software companies and pointed it at one thing: the firm’s AI capability, in whatever form it currently takes. Today the most-shipped form is <em>skills</em>: markdown files that encode a consulting workflow. A year from now <em>skills</em> will probably be a deprecated word, bundled into plugins or whatever replaces them. The weekly cadence is the constant; the artefacts are variable.</p>

<p>accilium consultants sort themselves into four classes: Rookie, Apprentice, Ninja, Pioneer. Rookies consume. Apprentices use a skill on real client work and give feedback, but do not build. Ninjas test what is not yet ready for everyone. Pioneers build. We have roughly ten Pioneers in a firm of about a hundred.</p>

<p>That distribution is the point, not a phase we are growing out of. The model we borrowed from the AI labs is not “every consultant a builder.” It is a thin builder tier producing artefacts that fan out through a reviewer tier to a mass of users. Open-source projects work this way: most committers are not contributors, most contributors are not users, and the counts go up by an order of magnitude at each step down. A skill written by one Pioneer this Friday is, a few weeks later, how an Apprentice in Bucharest handles an RFP without thinking about it. That sentence is the whole product.</p>

<p>Underneath, this is skill lifecycle management: a skill is in-progress, in preview, or live, and promotion between those is metadata, not engineering.</p>

<p>There are three problems we have not solved. The first is pace: new tools arrive every fortnight, and consultants - like everyone else right now - feel the anxiety of things moving faster than they can absorb. The second is maturity: we are still learning what holds up, and some of what we ship breaks in real use before we catch it. The third is commercial: transforming the firm internally is one thing; helping clients run the same transformation is another, and the offer is still too bleeding-edge for most clients to buy.</p>

<p>The lesson, eighteen months in, is that the hard part is not the technology. The tools have got good and they keep getting better, mostly without our help. The hard part is the firm: its sequence, who builds and who consumes, what gets rewarded, what gets quietly tolerated. A consultancy is a particular kind of organisation, with particular reflexes about expertise and authority and what counts as work. AI presses on every one of those. We did not choose this; we chose to take the pressure seriously and let it change how the firm is built. The result is not finished. It will not be finished. That is the point.</p>]]></content><author><name>Peter</name></author><summary type="html"><![CDATA[accilium is not becoming a software company. We are still consultants; the unit of value is still a decision made well, written down, and delivered on a slide. What is changing is what an organization is. A consultancy used to be a collective of people held together by hierarchy, rules, and shared method. AI is turning organizations into something closer to intelligent systems, where capability lives partly in people, partly in the tools they share, and partly in the cadence that connects the two. The operating model has to evolve with that. How we evolve the firm is itself becoming part of the work.]]></summary></entry><entry><title type="html">Welcome to aIQ</title><link href="https://accilium.github.io/aIQ-blog/2026/05/12/welcome-to-aiq/" rel="alternate" type="text/html" title="Welcome to aIQ" /><published>2026-05-12T00:00:00+00:00</published><updated>2026-05-12T00:00:00+00:00</updated><id>https://accilium.github.io/aIQ-blog/2026/05/12/welcome-to-aiq</id><content type="html" xml:base="https://accilium.github.io/aIQ-blog/2026/05/12/welcome-to-aiq/"><![CDATA[<p>Welcome to <strong>aIQ</strong>, a blog about accilium’s artificial intelligence journey — the ideas behind it,
the systems we build, and the practical lessons we pick up along the way.</p>

<h2 id="what-youll-find-here">What you’ll find here</h2>

<ul>
  <li><strong>Deep dives</strong> into models, papers, and techniques worth understanding.</li>
  <li><strong>Engineering notes</strong> from building real AI products: evals, prompts, infra.</li>
  <li><strong>Short takes</strong> on what’s changing at accilium as we ship new builds week to week.</li>
</ul>

<h2 id="why-another-ai-blog">Why another AI blog?</h2>

<p>We are tired of all the writing on LinkedIn and want to share insights on our journey in a way that feels right for us.
GitHub and GitHub Pages feels right, as most of the thinking we produce is written in markdown and managed in Git. We are also committed to open-sourcing skills and assets we create in public repositories to inspire others to follow us.</p>

<h2 id="how-this-blog-works">How this blog works</h2>

<p>Posts appear roughly every two weeks. Each one is signed by its author — an individual perspective, not a corporate position from accilium.</p>

<p>Stay tuned — more posts coming soon.</p>]]></content><author><name>Peter</name></author><summary type="html"><![CDATA[Welcome to aIQ, a blog about accilium’s artificial intelligence journey — the ideas behind it, the systems we build, and the practical lessons we pick up along the way.]]></summary></entry></feed>