It may seem odd to begin the discussion of content and context with data, but stay with me. The importance of having good data hygiene, good reference data sets, and data standards for interoperability is growing as we need to present better visualisations or crunch data in more sophisticated ways. The ability to carry out increasingly complex data analysis often means bringing multiple data sets together, using more sophisticated algorithms, to perform more complex calculations.
However, data does not tell the whole story; presenting that data doesn’t mean that the audience will necessarily understand what the data means. The role of context is provided by content. In a 2018 study discussed in Science Daily, between half and three-quarters of the study group had trouble interpreting statistical data, particularly when presented as probabilities instead of natural frequencies. Content fills the gap by adding context. That context completes the story by filling in the gaps of “what does that data actually mean?” to the audience segment who are not data scientists or statistics enthusiasts or not experts in a particular field.
Data and content: complementary elements
For all the importance of content to be delivered along with data, content has been largely ignored in the data arena. The manipulation of content to turn copy into meaningful context bears little resemblance to the processes for bringing meaning to data. Editing data means changing and possibly normalising a data point in a database cell. Editing content means checking the accuracy, the consistency of the language, the spelling and grammar, and most importantly, the context. Sandwiching a sentence between two other sentences is not a neutral act; it can enhance the telling of the data story or backfire terribly and not only detract from the story but offend your audiences in the process.
Salad and seasoning
A database designer once gave me a good analogy for how she viewed content and data, and I’ve adopted it as my go-to explanation.
Data is like salt and pepper. When someone wants salt, they can shake out a few or many grains of salt. All grains from the salt shaker are valid, as long as the salt shaker is filled with salt; all grains of pepper from the pepper shaker are valid, as long as the pepper shaker contains pepper.
Content is like a salad. Some ingredients are torn (such as lettuce leaves). Some are sliced (cucumbers). Some are diced (avocado). Some are grated (carrots). And some are left whole (cherry tomatoes). This variety of shapes, sizes, textures, and flavours is what adds to its visual appeal. Using a one-size-fits-all shaker doesn’t work.
I thought about this a lot as I prepared a series of salads last weekend for a party. I couldn’t bulk prepare anything because the salads were so different: cabbage and carrot slaw with a satay dressing, a watermelon and feta salad, potato salad with my late mother’s signature dressing, a vegan fusilli salad with vegetables and chickpeas, and strawberries with aged balsamic vinegar and black pepper.
Here’s where things get interesting. We can consider content as the salad, and salt and pepper as the seasoning. Data gets sprinkled in to provide specific information, but without the surrounding content, the data would lack enough context to be understood. In the world of communicating information, here are some examples of how it looks in real life - these are random examples taken from AI generated overviews on the internet.
Barclays UK personal savings rates feature up to [datum]% AER variable on the Rainy Day Saver (up to £[datum]).
Canada’s nominal Gross Domestic Product (GDP) is approximately [datum], ranking as the [datum] largest economy in the world.
In the following FIFA World Cup example, the data could be the scores, or the scores and country names and flags; the rest is content.
There will obviously be examples where data sets do the heavy lifting, with little surrounding content, but in a typical day, the average person is more likely to interact with information similar to the examples above: content that is “seasoned” with data.
Different entities, different needs
Management of content is a very separate discipline from management of data, and although the automated delivery of content dovetails with the need for automated data delivery, that’s where the similarity ends. And I’m not going to complicate this with the discussion of digital asset management, which has its own unique properties and complications.
When teaching, I define content as human-usable, contextualised data. My go-to example as follows. If I give you a data point of “12” and ask what it means, you can’t really tell me. If I add context, say, “it’s a month”, then you can use that to infer the month of December. That’s content. When you add more context, such as “December is when we have time off over the Christmas holidays”, and you have information. Add more context, such as “if you want to travel over the holidays in December, you’d better book early”, and you have knowledge. (I don’t think wisdom can be codified, personally. For example, look at the wide divide in political opinion with the same sources of information.)
There is a very strong temptation to treat content like data, confining content to cells in a database, moving it around in static chunks like so much boxed cargo. The mechanisms meant for processing data simultaneously limit content in so many ways. The complexity and nuance are dampened; the contexts are limited; its potential is hobbled. The editing process becomes cumbersome and error-prone; content bloat occurs as copies are pasted into multiple database cells; and the overhead of content maintenance becomes unwieldy.
To allow content to operate at its full potential, it needs to use its own standards and its own semantics, which ultimately enables its ability to interoperate. I’ll cover some of the issues of enabling content interoperability in more depth in another article.
Content as the training ground for AI
I sometimes start a presentation by asking “have you read the data, Moby Dick?” or “have you read the data called the Bible?” to illustrate the disconnect between what we call content and what we call data. The big controversy about AI being trained on copyrighted material isn’t about numbers in a database; it’s about sentences, paragraphs, and chapters. It’s about language and meaning and nuance.
Those who reduce content to “data” and data to “raw data” are doing content a great disservice. It’s akin to calling an F1 race car a “bicycle” and a bicycle a “basic bike” because they are both modes of driver-propelled transportation. I’ve watched data scientists try to use reference data sets to do things better suited to a content lifecycle because “how many changes could there possibly be" (many, it turns out), and then complain that the UX writers missed instances of those changes across multiple data sets. I’ve seen entire paragraphs shoved into a database cells, making content audits long, expensive, and error-prone. I’ve seen…well, let’s not go that route; the examples are sordid and many.
There’s so much more I could say about this, which I’ll break into logical topics and cover in separate articles. But for the love of all things semantic, can we stop collapsing two distinct entities into one sloppy bucket labelled data?
Note: A variation of this article was published in 2021, and has been updated to reflect changes in the industry.




