Writing

The cost of what you can't find

Material gets stored and backed up. Being able to retrieve one piece of it is a separate job, and it rarely gets funded.

Sep 2026·8 min read

An index kept by hand. Publishers were still paying people to do the same work seventy years later. Library of Congress, FSA/OWI Collection.

About twelve years ago I was in London working on a publishing platform for Hindawi, an academic publisher. I was a business analyst, working out what the solution needed to do and then building it with the team. The work was a content staging system in Drupal, wired into what they called their bank of stories, so a story could move from submission through review and out to publication without being handled twice. Publishing time on that pipeline came down from roughly fifteen days to about four hours.

Their people tagged every story by hand before any of it reached what I had built. Subject, type, status, the relationships between stories. It took considerably longer than the code did, and I remember finding that mildly inconvenient rather than interesting. It did not occur to me for years that the tagging was the part making any of it work.

What made it work

Every story arriving at that platform had already been described. Hindawi did the describing by hand, because at the time there was no easier way to do it.

Most organisations holding large bodies of material have never done it at all. A law firm's decade of matters, a bank's research, an insurer's claims files, any broadcaster's tape library: the material is stored, backed up and catalogued after a fashion, none of which is the same as being able to find one specific thing when somebody needs it.

Pharmaceutical field teams

The most developed version of this I have come across is in pharma.

The structure is a trigger, an action and a reason. Something fires: a formulary change at the account, a prescriber who has lapsed, a medical inquiry still open from a previous visit, an objection raised last time and never closed. Against that trigger the system surfaces the one action most likely to matter. It attaches the reason, because a representative deciding what to raise in the next ninety seconds will not act on a recommendation they cannot account for.

What may be surfaced is bounded by what may be said. Off-label material, unapproved claims, anything outside the licensed indication: none of it reaches the screen no matter how relevant it is. So it is not really search. It is retrieval working inside a permission boundary, and the boundary has to be part of the design rather than a filter added at the end.

None of it works unless the material was described and classified first. Pharma has gone further with that than most sectors, largely because regulators require it. Elsewhere the same precondition holds without a regulator enforcing it, and the cost of skipping it is highest where the material is not supporting the product but is the product.

Broadcast and publishing archives

Take broadcasters and publishers today, who hold decades of recorded material. A good deal of its remaining commercial life sits in reuse: syndication, licensing, building new work out of old, answering a question this year with something recorded twenty years ago. All of which depends on being able to locate a specific piece of it, say what is in it, and confirm the rights allow it to be used again.

The request that comes in looks like this. A production company in another market wants twenty or thirty seconds of a particular person speaking about a particular subject. They know roughly when it went out. They need an answer inside a week, because after that their edit is locked and they will use something else.

The footage is somewhere in the library. For the broadcaster to license it to them, three things have to happen inside that week.

01 Somebody has to find the clip.
02 Somebody has to confirm what was said and who said it.
03 Somebody has to establish that the rights permit licensing it into that territory.

If any one of the three cannot be done inside the week, the reply that goes back says the footage cannot be confirmed in time, which is a courteous way of declining.

The requester goes elsewhere or drops the segment. Because the request was declined rather than lost, it does not appear in any revenue report, and the material involved was paid for long ago and has been costing money to store since. Whatever the fee would have been, most of it was margin.

What enrichment produces

The missing part is the context that was never captured alongside the recording, because at the time everyone in the room already knew it.

Ask what enriching an archive involves and most people describe the content: a transcript, who is speaking, what is on screen. That layer is necessary and it is the one everyone pictures.

It is not what decides whether a second of the material can earn anything. That is the rights window, the territories permitted, whether the talent consented, whether an embargo applies and when it lifts. Clearing is the step that most often fails, and it is rarely filed under metadata. It lives with legal, in contracts and correspondence rather than anything the archive can query, so it sits outside most budgets for describing material. The information that gates the money tends to be the information least accounted for.

With both layers in place an archive can answer the request as asked: find me twenty seconds of this person on this topic that we can license in this territory.

A model on top of an un-enriched archive doesn't make it findable. It makes the un-findability faster.

What else the same gap blocks

A declined request is the smallest version of this, and the only one that leaves a trace.

The same records decide larger questions. Catalogue depth is inventory for anyone distributing at volume, so material that cannot be surfaced cannot be put into a channel whatever its quality. And where rights cannot be established, the material is difficult to offer to anyone, difficult to withhold on any firm basis, and difficult to pursue if it turns up somewhere it should not have.

Those are three separate decisions running on the same records. An organisation that cannot say what it holds and on what terms is not well placed to take a position on any of them.

Why it never gets funded

Enrichment has no demonstration. What it produces is only visible through something else working afterwards, and anything with a visible output competes better for the same money.

It has no single owner either. Syndication wants to fulfil requests, production wants to reuse footage, and legal wants to know what can be cleared. Each of them can point at a result of their own. Enrichment supports all three, which in practice means none of them puts it in a budget.

The harder problem is that its value depends on requests that have not arrived yet, so there is no way to forecast it. The case for enrichment gets made in conditionals: if the material were described, these requests could be answered. The cases it competes against get made by demonstration. A syndication lead who can show a signed deal this quarter is making the better-evidenced argument, and they are right to make it that way.

Work gets funded inside an organisation when it has a name, a number attached to it, and someone who notices whether it worked. Enrichment usually arrives with all three missing.

What has changed

The tagging Hindawi did by hand is now largely machine work. Speech to text. Speaker identification. Entity and topic extraction. Footage tagging. First pass flags for rights and personal information.

That is an observation about cost rather than a recommendation. The thing that made this uneconomic has changed, and whether it has changed enough to alter any particular calculation is for whoever holds the archive to work out.

Two qualifications keep it accurate. Machines do the labour now, and the judgment still needs somebody who knows the material: which questions matter, which decisions the archive has to serve, how rights get cleared, what the taxonomy should be, where provenance has to hold.

It is also a different use of the technology than producing content. Both run on models, but what enrichment produces is metadata for a retrieval system to work against, not something written for a person to read.

When this does not apply

Where material is held for compliance or simply kept, and no one is trying to retrieve anything from it, storage is a reasonable answer on its own. Material with no reuse value, no licensing interest and no bearing on a decision does not become more valuable once it has been described.

Why you probably are not doing this

If none of this has been done at your organisation, the reasons are likely to be sound ones. The requests that failed were each handled by someone doing their job properly, and a declined request is a closed matter rather than a problem anyone was asked to escalate. The losses are real but they are distributed across a year, a handful of people and several inboxes, so no one is looking at them together. Meanwhile the material itself is safe, backed up and cared for, which makes it genuinely difficult to see that anything is wrong.

Add to that a decade in which describing an archive properly cost more than the reuse income could justify. Declining to fund it was the correct decision for most of that period, and the habit of treating it as unaffordable has outlived the arithmetic that made it true.

Why it is still worth a look

The check is small. Pull the requests your organisation turned down over the last year, from syndication, licensing and production. For each one, note which of the three checks failed: the material could not be found, its content could not be confirmed, or its rights could not be cleared in time. Then note what the fee would have been.

That gives you a number rather than an argument, and it tells you which part of the problem you actually have. An archive failing mostly on rights needs different work from one failing mostly on search, and the difference decides where any money should go first.

It also bounds the job. You are no longer describing an archive, you are describing the parts of it people keep asking for, which is a far smaller thing and the only part with demand already proven. Organisations that try to describe everything about everything tend to produce a catalogue that goes unqueried.

And it supplies the three things that were missing. The activity gives the work a name, the declined requests give it a number, and the people who sent those replies will know whether it worked. If the number comes back small, you have your answer and it cost you a week. If it does not, you have the first case for this work that does not rest on a forecast.

What you actually own

Holding a library and being able to use it have come apart, and few organisations have measured the distance between the two.

On paper the archive is an asset: tapes, masters, rushes, interviews, decades of finished programmes. In practice what you own is the part you can describe, because every use of it, commercial or defensive, depends on being able to say what is in a recording and on what terms you hold it. The rest is footage you are paying to keep.

← All writing Work →