Nobody Was Building for AI
#261006 Product

Nobody Was Building for AI

, 3 min read

Why AI-first usually means data-first, told through a news CMS, a B2B sales database and an automotive publisher.

At Quintype, every news story was built of different elements. I led product on Bold, a headless CMS for news publishers. A conventional CMS stores a story as one block of formatted text. Bold stored it as a sequence of typed elements, so a paragraph, a quote, an image and an embed were each a separate piece that could be addressed on its own.

We did this for reasons that had nothing to do with AI. A headless CMS has to serve the same story to a website, an app and whatever surface comes next, and it can only do that if it knows what each piece of the story is.

Years after I left, Quintype moved into AI for the newsroom. I was not part of that work, and I am sure the team had plenty of its own problems to get through. Restructuring the content was not one of them.

The same was true of the archive. When content is stored as typed elements, every story a publisher has ever filed is usable by a model on the same day as the newest one. A publisher whose archive sits in page blobs has to pay for that conversion before any AI feature can touch it. If I were asked to design a CMS for AI today, I would build the same structure we built then.

I have since seen the same thing at two other companies. The phrase "AI-first" suggests the model is the starting point. In my experience the model is the last part to arrive and the easiest to swap. What takes years is the data underneath it, and that was usually structured by someone working on a different problem.

ZoomInfo and the companies we did not score

ZoomInfo is a company built on structured B2B data, which is why Account Fit Score could exist at all. AFS scored accounts against a customer's own closed-won deals in place of a static ideal customer profile. It was classical machine learning with no LLM involved, and the data science team built the model. Most of my time as the product owner went into decisions about the data around it.

The clearest one was scope. It was tempting to say we scored every company in the database. A score is only useful if someone can act on it, and a sales team cannot act on a company with no contacts attached. Only about a third of the companies had at least one contact. We scored those, which cut scoring compute by roughly two thirds. ZoomInfo's data had been built to be sold to sales teams, which is why the contacts were there to score against.

Cartoq and the car as a record

Cartoq is an automotive publisher in India that I consult with. Their content lived in WordPress, where a car review is a post with a title, a body and some tags. Everything the publication knew about a car's price, engine and variants sat inside paragraphs, readable by a person and opaque to software.

We moved the content to Strapi and modelled cars as structured records. Once a car is a record with fields, comparing two cars is a query, and the publisher starts to look like a car data company that also writes articles. Cartoq had come with a different question, which was why their content was hard to find and hard to sell against.

What I ask before funding an AI feature

These cases left me with four questions that I now ask before an AI feature gets a place on a roadmap.

  1. Can each piece of the data be addressed on its own? If the unit of storage is a page or a blob, the first project is restructuring, whatever the roadmap calls it.
  2. Who filled in the data, and was it a choice or a default? Volume says very little. A field that most records share because nobody changed the default tells a model nothing, and the model will use it anyway.
  3. Can someone act on the output? Scoring every company in the database makes a better slide than scoring a third of it.
  4. Does the structure cover the archive, or only what gets created from now on? For publishers and for most B2B data companies, the history is where the value is.

When the answers are poor, the honest roadmap item is a data modelling project, and it should be named and funded as one. It will not demo well, and it will take longer than the feature that sits on top of it. I would still fund it first, because the feature built without it produces answers that look right and cannot be trusted.

Models will keep changing every few months, and most teams will be able to rent the same ones. The data a company has structured, and the care with which it was filled in, is the part a competitor cannot rent. At Quintype it began as a way to get one story onto a website and an app.