Our Secret Sauce Is Still Human

I’m camping this week, so rather than writing another long post, I thought I’d keep this one short and talk about something we don’t talk about enough at ListEdTech: How do we actually get our data?

With AI everywhere, it’s tempting to think that collecting data is mostly about having the right crawler, AI model, or algorithm.

Yes, we do use technology. We have different types of crawlers and automated processes that help us identify signals that an institution may be using a particular product. We monitor thousands of public sources, including university and company websites, press releases, RFPs, contracts, board documents and social media. Our systems process more than 25,000 new alerts every day.

But that’s only the beginning.

Our secret sauce is humans. Every new data point that goes into our database is validated by a human. Why? Because identifying that a product is mentioned somewhere online isn’t the same thing as knowing that an institution actually uses that product.

For example, a university website might mention a software company because:

  • They are evaluating the product.
  • They issued an RFP.
  • The product is used by one school, but not the entire university.
  • The product is being replaced.
  • The vendor is a campus partner or a conference sponsor.
  • The information is simply outdated.

A crawler can find the mention. A human needs to understand the context.

We don’t pretend that we’re perfect, we make mistakes. Any company working with a dataset this large will. The important thing is having a process to find those mistakes and correct them. Every month, our data quality team audits more than 3,000 database entries. We also conduct an annual review to confirm that institutions are still using the systems we’ve identified.

Today, our database contains hundreds of thousands of product implementations across 80 product categories and 81,000+ institutions worldwide. That’s a lot of data. And it would be impossible for humans to collect all of it manually. But it would also be a mistake to assume that AI can simply collect it all for us. The best approach is to use both in a rigorous process.

Machines are good at finding signals, and humans are incredibly good at understanding context. That’s how we approach our data at ListEdTech. Scale comes from automation; trust comes from human validation. With hundreds of institutions using our portal every month to help improve our data, that collaborative loop is what keeps our information sharp. When you’re making strategic decisions about a market, competitors, customers or technology adoption, I think that distinction matters.

You can learn more about how we collect and validate our data here: ListEdTech Data Overview.

Latest Posts from ListEdTech