AI agents for the work someone has to read first.
We build AI automation and AI agents for specific, repeated tasks — reading documents, routing messages, drafting replies — with a human check wherever the result matters. AI is one option, not the starting point: when a deterministic workflow solves the task, that is what we build.
Let's talkWhere this usually starts
Invoices, delivery notes and long email threads pile up because someone has to read each one before anything can happen. The work is not the decision — it is the reading.
An AI demo impressed everyone six months ago. It is still a demo. Nobody could say how often it was wrong, so nobody dared put it in front of customers.
Your most experienced people spend their days answering questions whose answers already sit in your own documents — contracts, manuals, past quotes. Nobody else can find them there.
There is pressure from above to do something with AI, but no one has named the specific task it should do. That is exactly how pilots get built and shelved.
What you actually get
The use case, chosen honestly
We start from the task, not the technology. If a deterministic workflow solves it — fixed rules, same result every time — we tell you and build that instead. AI goes where the input genuinely varies: language, documents, judgment.
Pipelines that read for you
Invoices, delivery notes and scans run through OCR, and the extracted fields land in your systems. Anything below a confidence threshold goes to a person instead of being guessed at.
Agents with a narrow job
An AI agent reads the situation first, then picks its next step from a small set of tools we define — look up an order, draft a reply, hand over to a person. Connected over MCP, with no way to act outside that set.
Answers that cite their source
Internal Q&A grounded in your own documents through retrieval over Supabase and pgvector. Every answer points to the passage it came from, so checking it takes seconds — and an answer without a source does not go out.
Approval steps, built in from day one
Drafted replies wait in a queue for a person to approve. Extracted data can be corrected before it reaches the ledger. The human check is part of the design, not a patch after the first incident.
Evaluation before and after go-live
We build a test set from your real cases and measure accuracy before launch, then log every call in production. You get a number for how often the system is right — not a feeling.
How an engagement runs
- 01
One task, and a definition of correct
We pick a single reading-heavy task and write down what a correct result looks like, who verifies it, and what happens when it is wrong. If correct cannot be defined, we do not build.
- 02
Measured on real examples
We collect past cases — real invoices, real threads, real questions — and test until the accuracy is known. That number decides whether the project continues, and you hear it either way.
- 03
Live, with a person in the loop
At launch, every output passes through human approval. We loosen the checks only where weeks of logs justify it — and some steps keep them permanently, on purpose.
What we build with
- Claude
- OpenAI
- MCP (Model Context Protocol)
- n8n as orchestration
- Supabase + pgvector
- Document parsing & OCR
- ElevenLabs
- Perplexity
- Human-approval steps
- Evaluation & logging
Brain — a shared memory for a team's AI tools
A knowledge graph we built for our own work: notes, decisions and project context exposed over MCP, so every AI tool the team uses — in any CLI, in any session — starts with the same memory instead of a blank context window.
What teams say after working with Radek.
Real feedback from businesses using the systems, automations and websites we have delivered.
“The collaboration was very professional. I especially appreciate Radek's fast communication, willingness to respond to our needs and ability to find practical solutions even for specific requirements.”
“Everything works great, often even beyond the scope of the brief. I especially appreciate the speed and the patience to explain everything through meetings or clear video walkthroughs.”
“The result is a modern, clear website, technically well crafted and precisely matching our needs. Professional approach, reliability and speed that truly makes a difference.”
Questions about AI
What does AI do reliably today, and what does it not?
It reads well and it drafts well: classifying and routing messages, extracting fields from documents, summarising long threads, drafting replies, answering questions from a defined set of documents. What it cannot do reliably is guarantee a correct result without a check, or give the same answer to the same question every time. We build on the first list and design a human check for everything on the second.
How do you stop it from inventing things?
Three ways, used together. Answers are grounded in your own documents and must cite the passage they came from — no source, no answer. Outputs that matter pass through a person before they take effect. And everything is logged and measured against a test set, so a drop in accuracy shows up in numbers, not in a customer complaint.
Which models do you use, and why?
Mostly Claude and OpenAI models, chosen per task rather than per fashion. The criterion is boring: measured accuracy on your own examples, at a price that makes sense for the volume. Extraction jobs often run fine on small, cheap models; judgment-heavy steps get bigger ones. When models change, we re-test — because they do change.
Does our data train anyone's model, and where is it processed?
No. We use the commercial APIs, whose terms exclude training on your data, under a signed data processing agreement. Retrieval indexes and logs live in an EU-hosted database under your control. And where data must not leave your infrastructure at all, we design for that — it narrows the model choice, and we will tell you what that costs in quality.
What is an AI agent, and how is it different from ordinary automation?
An ordinary workflow follows steps you fixed in advance — same input, same path. An agent reads the situation first, then chooses its next step from tools it has been given: look something up, draft something, hand over to a person. That freedom is the benefit and the risk at once, so we keep the toolset small, the permissions narrow, and irreversible actions behind human approval.
When would you talk us out of AI?
Whenever the rule can be written down. If the task reads 'when X arrives, do Y', a deterministic workflow is cheaper, faster and never makes things up — and that is what we build for most of the automation on this site. AI earns its place only when the input is too varied for rules: free text, messy documents, judgment.
Related services
- n8n Automation & Workflow Developmentn8n automation for European SMBs — connect forms, CRM, email and databases so your team stops moving the same data between tools by hand.
- Make.com Automation & Integromat MigrationMake.com agency for European SMBs — Integromat migration, cleanup of overgrown scenarios, operations cost audits, and a straight answer on Make.com vs n8n.
Bring us one task, not an AI strategy.
Name one thing your team reads, sorts or answers over and over. We will tell you whether AI handles it reliably — and if a plain workflow does it better, that is what we will recommend.
Let's talk

