MICKAI®ArticlesMore Than Half Your Company Data …
Article · 11 August 2026

More Than Half Your Company Data Is Dark: Put the Assistant on Your Own Brain, Not the Cloud

Over half your data sits unread. An offline Assistant on your own hardware turns it into answers without a per token cloud bill.

Author
Micky Irons
Published
11 August 2026
Follow Micky Irons
LinkedInX
dark-datasovereign-aion-premise-aienterprise-searchcost-reduction
More Than Half Your Company Data Is Dark: Put the Assistant on Your Own Brain, Not the Cloud

More than half of everything your company stores is dark data: information you captured and paid to keep, then left unread in file shares, scanned PDFs and line of business systems. You turn it into answers by running the Assistant offline on your own brain, the model built on your own files, so retrieval happens on hardware you own with no per token cloud API charge and no record ever leaving the building.

Dark data in 2026: most of what your enterprise owns, unread

Industry reporting puts around 55 percent of enterprise data as dark in 2026, captured and stored but never used for a decision, and most of it is unstructured: contracts, emails, forms, images, call notes and scanned paper. The original Splunk dark data research coined the phrase, and later storage indices put the unused share of corporate data anywhere from a quarter to three quarters. The uncomfortable part is that the volume keeps growing, with unstructured files the fastest growing segment, while the share you actually read keeps shrinking. Every year you pay to hold more information you never look at.

That makes dark data a cost line, not only a governance risk. You are paying to store it, paying people to hunt through it, and paying a second time when you finally reach for a cloud tool to read it.

What dark data costs you today

The bill arrives in four places at once, and the reporting is blunt about the size: some estimates put annual dark data storage waste in the millions for large organisations, with a meaningful share of firms spending over a million a year on data they never manage or query.

  • Storage: cloud object storage charges a growing monthly fee per gigabyte for files no one queries, the dark data storage overhang.
  • Labour: staff re key figures from paper and PDFs by hand and spend hours searching across silos for a document that already exists.
  • Decisions: when the answer is locked in an unread file, the decision gets made without it, and the rework lands later.
  • Cloud AI: the moment you point a cloud model at those documents to read them, you pay per token on every query and per page on every ingest, and the invoice scales with your team, not your value.

How the Assistant reads your dark data on your own hardware

Two studios do the work, and a studio here means a ready made application for a business function that runs inside one system. Omni is the Assistant, the sovereign front door where one prompt is routed across the offline brains on hardware you own. The Document studio handles long form writing and retrieval augmented drafting over the sovereign vault, the governed store of your own files. Together they read your dark data in place, with nothing shipped to a third party.

The mechanism is deliberately plain.

  • Point the sovereign vault at the file shares, scanned PDFs and line of business exports that hold your dark data.
  • The offline brains extract, index and structure the content on device, so unstructured paper becomes searchable without a cloud OCR fee.
  • Ask the Assistant in plain English, and it routes the question across the fifty brains and retrieves the grounded passage from your own corpus.
  • The Document studio drafts the answer, brief or report with the source passages attached, so the output is traceable.
  • Every retrieval and draft is sealed to the Open Audit Record on hardware you own, so you can show what was read and why.

Because the Assistant runs on the company's own brain built on its own data, the retrieval never leaves the building and there is no per token inference charge as usage grows. The dark data is read where it already lives.

What you replace, and what you save

What you run todayWhat it costs youWith Mickai
Cloud object storage for files no one queriesA growing monthly fee per gigabyte for dark dataFiles stay on your own disks inside the owned system, no per gigabyte cloud bill
Cloud LLM and RAG APIs to read documentsPer token on every query, multiplied across a teamAssistant runs offline on your own brain, no per token inference charge
Document intelligence and OCR services (Azure AI Document Intelligence, AWS Textract, Google Document AI, ABBYY)Per page on every ingestOn device extraction, no per page cloud fee
Enterprise search and copilots (Glean, Microsoft 365 Copilot)Per seat, per user, every monthOne owned system, no per seat search subscription
Manual re keying and document huntingStaff hours re entering paper and PDFs by handRetrieval and drafting handled in the Document studio, hours returned to the work
Shipping records to a third party processorVendor risk and cross border transfer exposureNothing leaves the building, no third party processor

The evidence trail, sealed on device

Reading dark data usually means moving sensitive records to a cloud processor, which is exactly what a regulated business cannot afford to do casually. Here the data never moves. The system is sovereign and on device: every AI action is sealed under post quantum cryptography into a signed record, the Open Audit Record, so you can show which document was read, by which brain, to produce which answer. That evidence supports SOC 2, ISO 27001 and GDPR examinations rather than replacing them, and it is produced as a by product of normal work rather than assembled by hand before an audit.

Where the money goes instead

Consolidating dark data onto an owned system turns three recurring meters off at once: the per gigabyte fee for storing files no one reads, the per token and per page cloud AI charge for reading them, and the per seat search or copilot subscription layered on top. The spend moves from an invoice that grows every month into hardware you own and keep. The labour that used to go into re keying and hunting for documents goes back to the work the documents were for.

The paper does not disappear, but the tax on ignoring it does.

Frequently asked questions

What counts as dark data?

Dark data is information your organisation collects and stores but never uses, from scanned contracts and forms to emails, call notes, images and line of business exports. Industry reporting puts it at roughly 55 percent of enterprise data in 2026, and 80 to 90 percent of enterprise data is unstructured, which is where most of the dark data sits.

Does reading it with the Assistant send our documents to the cloud?

No. The Assistant runs offline on your own brain, built on your own files, and the Document studio retrieves over a sovereign vault held on hardware you own. Nothing is shipped to a third party and there is no per token cloud API charge.

What cloud costs actually go away?

The per gigabyte fee for storing data you never query, the per token and per page charges for cloud models and document intelligence services reading it, and the per seat cost of an enterprise search or copilot subscription. They are replaced by one system on hardware you own.

Can we prove what the Assistant read for an audit?

Yes. Every retrieval and draft is sealed to the Open Audit Record on device, so you can show the source document, the brain that read it and the answer produced. That evidence supports SOC 2, ISO 27001 and GDPR examinations rather than claiming a certification you do not hold.

Subscribe
Get every new Mickai article by email.

Long-form essays on sovereign AI from Micky Irons. One email per article. No tracking, no marketing, no third parties. Every email includes a one-click unsubscribe link.

Prefer RSS? Subscribe at /articles/feed.xml.

Originally published at https://mickai.co.uk/articles/dark-data-assistant-on-your-own-brain. If you operate in a regulated sector or want sovereign AI on your own hardware, the audit form on mickai.co.uk is the entry point.
More articles
18 Aug 2026
How Telecoms Operators Meet the Telecommunications Security Act With AI That Never Leaves the Network
Telecoms operators meet the Telecommunications Security Act code of practice with AI that runs inside the security-critical boundary on operator-owned hardware. A zero-egress perimeter keeps network configuration and signalling data within operator control, so nothing sensitive crosses out to a public cloud service.
18 Aug 2026
Can energy operators run AI on grid and OT data on-premise to satisfy the Cyber Assessment Framework?
Yes. Energy operators can run forecasting and anomaly detection on grid and OT data entirely on their own hardware, and this satisfies the Cyber Assessment Framework more cleanly than cloud analytics, because telemetry never leaves the audited perimeter and no third-party processor exists to assess.
18 Aug 2026
How Airports Meet EASA Part-IS from February 2026 with On-Site AI
Part-IS applies to aerodrome operators from 22 February 2026 and makes the airport, not its vendor, accountable for information-security risk. Running AI on operator-owned hardware behind a zero-egress perimeter keeps passenger and operational data inside that boundary, so a supplier's SOC 2 cannot discharge it.
18 Aug 2026
Can Automotive Suppliers Use AI on OEM Design Data While Keeping TISAX Prototype Protection?
Automotive suppliers can run AI on OEM design and prototype data and keep TISAX prototype protection, but only when the model runs on their own hardware inside the protected zone. Public cloud AI transmits the data outward, which prototype protection forbids.