Most of us have far more digital history than we realise.
It sits across current computers, old computers, external drives, NAS boxes, cloud accounts, backups, exported archives, forgotten folders and copies of copies — and so on. Some of it is obvious. Some of it has lost its context completely.
We tend to call all of this data. I think that is too simple.
What we actually accumulate is a digital estate.
A digital estate is not just the files themselves. It includes the relationships around them: where something came from, what it became, what it duplicates, what depends on it, who or what it represents, and the history that explains why it still exists.
That distinction really matters, because storing information and understanding it are very different problems.
The term digital estate is commonly used in estate planning and digital inheritance to describe somebody's digital assets, accounts and information after death. That is an important use of the term, but I use it more broadly.
To me, a digital estate is the accumulated body of information, systems, identities, relationships, copies, records, histories and dependencies that builds up around a person or organisation over time.
It exists while we are still very much alive and creating it.
Storage is not understanding
We have become remarkably good at keeping things.
Storage capacity has increased, disks have become relatively inexpensive, backups can run automatically, cloud services synchronise information between machines, and moving to a new computer no longer necessarily means leaving the previous one behind.
This is mostly a good thing.
The problem is that our ability to preserve information mechanically has grown much faster than our ability to preserve the explanation around it.
A filesystem can tell me that a file exists.
It can tell me its name, size, location and dates. I can calculate a hash and identify files that appear to contain exactly the same bytes.
Those are useful facts.
They do not necessarily tell me what the file means.
They do not tell me whether it is the authoritative copy, an obsolete copy, a backup, an export, an intermediate version, an attachment extracted from somewhere else, or something retained because nobody was quite certain whether it was safe to remove.
The storage is intact. The understanding may not be.
That is the distinction I find interesting.
Copies accumulate history
Consider something very ordinary: replacing a computer.
You copy your Documents folder to the new machine.
You keep the old computer for a while, just in case.
Perhaps it was also being backed up by Time Machine. A few years later, some of its contents move onto a NAS. Another part is synchronised through a cloud service. At some point you export material from an account before closing it.
Eventually, you may have several apparent copies of the same history.
Which ones are redundant?
That sounds like a storage question, but it often isn't.
One copy may represent the state of the old machine immediately before it was retired. Another may contain material added during a later migration. A backup may preserve files that were subsequently deleted. An exported archive may contain information that looks duplicated but includes metadata or relationships that were lost elsewhere.
Even byte-identical copies are not always historically identical.
Their contents may be the same while their provenance is different.
That does not mean every copy must be kept forever. It means that copy and redundant are not automatically the same thing.
Before deciding what something is, it helps to know how it got there.
Context decays
This becomes harder with time because context has a habit of disappearing before the information itself does.
A folder called Accounts 2017 probably made perfect sense in 2017.
A decade later, it might refer to financial accounts, user accounts, a client called Accounts, an exported database, or something entirely different.
A directory called OLD is even less helpful.
So are Backup, Archive, Final, Final2 and the inevitable Final Final.
The people who created those names understood them at the time. Often that person was us.
Then the machine changes.
The application that created the information disappears.
A website closes.
An account is exported.
A disk is copied.
Metadata changes.
A folder is moved out of the environment that gave it meaning.
The information survives, but part of its explanation does not.
This is one reason old digital material can become surprisingly difficult to reason about. Nothing is necessarily corrupted. Nothing dramatic has happened. The bytes are still there.
What has decayed is the context.
A photograph is a simple example.
IMG_4382.jpg tells us very little.
But perhaps that photograph was imported from an iPhone into a photo library, edited on a Mac, included in a backup, exported several years later, copied onto a NAS and then included in another archive.
The JPEG is only one part of that story.
Where it came from, what happened to it and how the surviving copies relate to one another are also information.
In other words, a digital estate contains relationships as well as objects.
“Can I delete this?”
This is where a seemingly mundane question becomes much more interesting.
Can I delete this?
Sometimes the answer is obvious.
Sometimes it is not.
Imagine finding five directories:
- Documents
- Documents Old
- Documents Backup
- Documents 2018
- Documents from MacBook
The simple approach is to compare their sizes, look for duplicate files and remove whatever appears unnecessary.
That might work.
It might also destroy the only copy of something that existed during one particular stage of a migration.
Before deleting safely, we may need to know whether one directory replaced another, whether their contents were ever merged, whether files disappeared between copies, whether one is a complete snapshot, and whether later material exists somewhere else.
That is not really a disk-space problem.
It is an evidence problem.
The difficult part is not pressing Delete. It is establishing enough confidence to know what Delete means.
This is why I think deletion, in a mature digital estate, is often a reasoning problem.
Cleanup and stewardship are not the same thing
There is an understandable tendency to approach old data as clutter.
Find duplicates. Sort the folders. Delete the rubbish. Free some space.
Sometimes that is exactly what needs to happen.
But cleanup begins with the assumption that the primary task is removal or organisation.
Stewardship begins with a different set of questions.
What is this?
Why does it exist?
Where did it come from?
What is it related to?
Does anything depend on it?
Is this the best surviving representation of something?
Should it be preserved, consolidated, documented, moved, or deleted?
Deletion is still one possible outcome. Often it is the correct one.
The difference is that it becomes a conclusion rather than the starting assumption.
That distinction matters more as a digital estate gets older.
A folder accumulated over six months is usually manageable by memory.
A body of information accumulated over twenty years, several computers, storage systems, online services and operating systems is something else entirely.
By then, you are not merely organising files.
You are dealing with history.
Why I became interested in this
I have been using computers long enough for my own information to have passed through generations of machines, disks, operating systems, backup arrangements and storage strategies.
Some material has migrated cleanly.
Some has been copied more than once.
Some survives because, at some point, deleting it seemed riskier than keeping it.
That last category is particularly interesting.
For an individual, keeping something because its significance is uncertain can be perfectly rational in isolation. The immediate cost of keeping another copy can seem trivial compared with accidentally destroying something important.
Repeat that decision often enough, across enough years, and you eventually create another problem.
You possess the information, but you no longer necessarily understand the estate it belongs to.
I became increasingly interested in what it would take to reconstruct that understanding.
Not simply: What files do I have?
But:
What represents the same underlying thing?
Where did these copies originate?
Which systems produced them?
What changed between versions?
What relationships survive?
What is genuinely unique?
What can be retired with confidence?
Those questions move the problem away from traditional file management and towards systems thinking, provenance and technical investigation.
That is the territory I find interesting.
From investigation to software
Eventually I started treating the problem less as a collection of messy folders and more as a system that could be investigated.
Cleverfeets is the software project I am building around that problem.
It is an attempt to understand a digital estate rather than merely enumerate what is stored in it: to examine identities, relationships, provenance, copies and history as part of the same problem.
The software is still in development, and the implementation will continue to change as I learn more about the problem.
But the underlying question exists independently of any particular software.
We are creating digital estates whether or not we call them that.
Every computer migration, backup, cloud account, archive, exported mailbox, old website, phone, hard drive and forgotten folder contributes another piece.
The significant question is not simply how much of it we can store.
It is whether, years later, we can still understand what we kept.
Information can survive longer than understanding.
A digital estate begins to make sense when we stop treating it as a pile of files and start treating it as accumulated history.
C. Matt Fletcher