Tax readiness
Am I actually ready to file?
The graph knows what evidence a tax year should produce and compares it to what has been observed. Readiness becomes a completeness score with the specific holes listed behind it.
A private AI architecture experiment: what becomes possible when personal data is reconciled across time, entities, relationships, and evidence.
local multimodal AI · entity resolution · temporal reasoning · human validation · knowledge graphs
A family's history is already written down: in three mailboxes, several drives, two photo libraries, and the dumps of every laptop that came before this one. Each account is indexed by a company for its own purposes, and none of them know the others exist. The information is not missing. It is unreconciled.
LifeOS reads those sources and builds one model above them: people, places, events, documents, obligations. Twenty files that are one document become one document with twenty pieces of provenance. The model is cross referenced, correctable by a human, and lives on hardware its subject owns.
A personal experiment on one family's archive, built in public. Two layers run today and are measured. The reasoning layer above them is design work, labelled that way everywhere on this page. Nothing here is a product.
LifeOS does not organize files. It organizes reality inferred from evidence.
Every layer takes the layer below it as raw material and refuses to discard what it cannot yet explain. Uncertainty travels upward with the data instead of being rounded off at each step.
Running today Being explored
Whatever the family already produced, in whatever state it was left in. Nothing is normalized before it is recorded, because the mess is data too.
Turning bytes into something with meaning attached. Text out of scans, coordinates out of photographs, faces into vectors, images into semantic space. Failures are recorded as failures.
The hardest layer, and the one everything above depends on. Deciding when many things are one thing, which copy is the real one, when two records describe the same person, and what to do when the evidence disagrees with itself.
The objects the system actually reasons about. Not folders and not filenames: entities with histories, connected to the evidence that supports them.
Every assertion carries a confidence and a trail back to what produced it. A human answer outranks anything inferred, permanently, and survives a full rebuild of the layers beneath it.
The layer this whole structure exists to make possible, and the one still being built. It is where the question a person asks their own archive changes shape.
Where is my document?
What should exist? What changed? What is inconsistent? What deserves attention?
The first question is search, and search has been solved for thirty years. The second set requires a system that holds a model of what a complete life record looks like, compares it against what it has actually observed, and can say something useful about the difference. That gap is the whole project.
Every line above has run against the real archive and has numbers behind it on this page or in the repositories.
Prototypes and design work, not scheduled runs. Several of these may turn out to be the wrong idea, and they are listed anyway, because a page that only lists its successes is not worth reading.
Choose any entity as the center: a person, a home, a policy, a tax year. The model rearranges around it, everything the evidence ties to it, at a distance set by the strength of the tie.
The same model arranged along time: eras widen into years, years into events, events into the evidence beneath them.
The model as a diff: each change since you last looked, carrying its receipts and a way to disagree.
Eight photo sources reconciled into one set. Checksums counted 56,943 unique files; perceptual hashing found roughly 45,325 photographs that actually exist. The collection was smaller, and more precious, than the byte count claimed.
The same tax document arriving as an attachment, saved to a drive, re-saved by a spouse, and scanned again years later is one document with four pieces of provenance, not four documents.
Identified, marked, reported, and left exactly where it was. Nothing in the pipeline has ever deleted an original, and no part of it is permitted to.
The graph that makes the rest possible. Typed links between documents and the world are what let a question about a person, a year or an institution be answered without a filename ever being involved.
Found across fifteen monitored record series. Nobody asked for those records. The system worked out what a complete series looks like and reported the holes, which is the first thing on this page that search cannot do.
Reconciles every photo library the family owns into one master set, then learns the people, places and events inside it, entirely on local hardware.
Perceptual hashing found the collection was smaller, and more precious, than the byte count claimed.
Catalogs every document the family has, works out which ones are the same document, extracts what they are about, and reasons about what is missing.
Extraction failures are recorded as failures. Absence is a conclusion, and it has to be earned.
Am I actually ready to file?
Waiting for evidence.
The mission does not close when a search returns.
It closes when the evidence is complete.
Am I actually ready to file?
The graph knows what evidence a tax year should produce and compares it to what has been observed. Readiness becomes a completeness score with the specific holes listed behind it.
What should still exist, but disappeared?
Recurring records are watched for breaks. A series that arrived every year for six years and then went quiet is a finding. This is absence reasoning, the opposite of search.
What might my history have left behind?
Evidence chains run from a former employer to a plan administrator to the absence of any rollover evidence, and end as a lead a human verifies with the institution. Here is the chain. Here is the gap in it.
If my family suddenly needed to understand everything important, could they?
A living map of what exists, who it belongs to, who depends on it, and what is missing, generated from evidence. The archive a family inherits should not be a password nobody wrote down.
What started but never concluded?
Applications without outcomes. Claims without resolutions. Deposits without refunds. The useful signal is the silence at the end of the chain, and silence is exactly what a search box cannot return.
Reconstruct the trip, not just the photos.
Photographs, the people in them, the coordinates, the reservation and the receipts collapse into one event with a beginning, an end, a cast and a cost. A gallery shows files from a week. This returns the week.
Written while building, in the order the lessons arrived. Every number below is a real measurement from this archive.
I merged similar face clusters with union-find, which quietly makes merging transitive: A resembles B, B resembles C, so A and C become one person. At six thousand face appearances this was invisible. At 53,780 it produced a single cluster of 16,504 and put sixty percent of the library into five clusters. Merging is now one hop, between mutually nearest clusters only.
Lesson: a similarity relation is not an equivalence relation, and treating it as one fails silently until it fails enormously.
The face toolkit ships an age estimator, so I used it, then measured it against four people whose ages I knew exactly. Wrong by 11 to 26 years in every case: a three-year-old read as forty-five. It was not demoted to a hint. It was deleted, and age now comes only from birth dates a human typed in.
Lesson: a measured-wrong signal shown as a hint is still a wrong signal, and it will be believed.
The review card showed the twelve sharpest crops from each cluster, which is exactly what you would do if you wanted the reviewer to say yes. A 1,650-face grab bag looked like one child, and it got confidently named. Cards now sample across the whole cluster, ugly crops included.
Lesson: if you are asking a human to catch errors, show them where errors would live.
6,884 duplicate clusters merge files with different bytes and identical pictures: a cloud-recompressed copy beside the original it was made from. And the most-copied image in the whole library, at 34 copies, was a user-interface icon from a code repository. Deciding what is not in the corpus is as much of the work as processing what is.
Lesson: identity is a property of content, and content is not bytes.
About 6,200 photographs carry dates recovered from folder names, and every one collapses to January 1. Fed to event detection, that is the largest New Year's party in history, every year, for a decade. They are grouped as an undated batch instead, because a date whose precision you invented is worse than no date.
Lesson: carry the precision of a value alongside the value, or something downstream will assume it.
A matcher hunting uncashed cheques flagged a religious text with high confidence, because "void after" appears in cheque boilerplate and, it turns out, in theology. Nothing was harmed; that analysis only ever produces leads for a human. But it is why no keyword hit is ever treated as a finding.
Lesson: precision is a discipline, not a default, and lead generation must always end at a person.
Nothing is stated without a path back to the evidence that produced it.
An answer a person gave wins permanently, and survives a full rebuild of everything derived.
The model is allowed to hold a question open. A confident guess is more expensive than a gap.
A failure to read is recorded as a failure to read, and never counted as an absence.
Duplicates are marked and left in place, and every rebuild carries human corrections forward.
Perception runs on the laptop at near-zero marginal cost. Frontier reasoning is spent only where it earns its price.
An AI that knows everything about a family, and nothing leaves the laptop.
The face, photo and document models are ONNX files on disk. There is no inference endpoint. Zero cloud calls are made to classify a photograph, embed a document, or recognise a face.
The review application binds to 127.0.0.1. It is not exposed to a
network interface, so there is no deployment in which it is accidentally reachable.
Mailbox access is read-only at the OAuth scope. The permission to send or delete was never granted, so the capability does not exist to misuse.
The source catalog is opened read-only at the SQLite driver. A write fails at the driver rather than at somebody's discipline.
This is the one architectural claim a cloud product cannot copy, and the reason the experiment is worth running. It is also why there is no login on this page: the application that shows the real photographs binds to the laptop and nowhere else. The public page is the story of the system, not a window into it.
Photo master set reconciled across eight sources, with the perceptual verdict: 6,884 duplicate clusters checksums could not see, one canonical copy per cluster, every member traceable.
Intelligence layer. Local faces, places, events, and the review loop whose answers outrank every model in the stack.
Document catalog and knowledge graph. 12,491 occurrences into 7,703 logical documents; 890 entities; 18,369 links.
Read-only mailbox ingestion, then the first absence reasoning. Three accounts scanned under a read-only scope, and the first two questions nobody asked: what is missing, and what value went dormant.
Cross-modal joins between the photo model and the document model, and the first Mission that keeps its state between runs.
What becomes possible when an AI does not merely know what you asked it, but gradually understands the structure, history, and unfinished business of your life?
LifeOS · ThirdOcular Labs
Built by Sandeep Kanuri