Digital, Marketing, CX

Does AI Actually Know Your Taste? I Scanned 12.8 Million Movies to Find Out

JJs Movie Muse AI Oracle vs Concierge

I scanned 12.8 million films and shows to teach an AI my exact taste in movies. It ended up predicting my ratings about as well as ignoring me completely.

The gap between my clever personal model and a dumb public average was 0.009 rating points. Statistically, that is a rounding error wearing a lab coat.

And yet I reach for it most evenings. The part that flopped taught me more than the part that worked.

Three years of chasing the wrong number

This is chapter five of a hobby that refuses to end.

Back in 2023, I had ChatGPT analyse 32,000 titles and my own ratings to build a personal recommendation engine. It lived in a spreadsheet. It was clever, and it kept suggesting films I’d already seen because it had no idea what I’d watched. In 2024, I turned it into a build-your-own guide you could run in minutes.

Every version rested on the same quiet assumption: a bigger, smarter catalogue would give me better picks. More data in, better evenings out.

So this year I went all the way. Using an agentic AI in work mode, I pulled IMDb’s full public datasets: 12.8 million titles, 15.6 million names, tens of millions of rows in all. I matched all ~900 films and shows I’d rated over the years, layered on the taste profile the AI had built from those ratings, and had it hand-screen a private catalogue of 4,004 titles I could pull up on my phone or in a browser.

Then, before I let myself get excited, I ran an honest test. That is where the plan came apart in the most interesting way.

The uncomfortable finding

I held back 170 of my ratings the model had never seen, froze its predictions, and only then checked the answers.

The scoreboard, in plain terms: predicting my rating from the public IMDb score plus a simple film-or-series adjustment got you almost all the way there. Adding my personal genre preferences and favourite directors on top improved the error by 0.009 points, and a proper confidence check said the real improvement could just as easily be zero.

No clear advantage. My own data, telling me my taste model was mostly theatre.

For a moment that stung. I’d scanned millions of records to learn something the crowd already knew. My ratings line up with the public average about 70% of the way, so “the internet’s opinion, lightly adjusted” was quietly doing my job for me the whole time.

Then I looked at what I was actually using the app for, and realised I’d measured the wrong thing entirely.

The Oracle and the Concierge

Here is the trap I’d fallen into, and I suspect a lot of the AI world is in it too.

I wanted an Oracle. A model that could look at a film and predict, to the decimal, how much I’d love it. That is the fantasy every recommendation engine sells, and it turns out to be both hard and, for me, nearly pointless. The crowd already predicts my ratings well enough.

What I needed was a Concierge. Something that handles the messy, unglamorous stretch between “I feel like watching something” and “we’re actually watching this.”

Picture a normal Friday. We pay for five streaming services. Twenty minutes in, we still haven’t chosen anything. Half of what the “smart” apps suggest I’ve already seen, and a good chunk isn’t even available where I live. That is the real problem, and no amount of taste-prediction genius solves it.

JJs Movie Muse

The Concierge version asks who’s watching first: just me, the two of us, or the whole family with the kids. It matches the mood and the time I actually have tonight. It never suggests something I’ve already rated. It tries to rule out anything I can’t watch where I am, on a service I already pay for, with the full season ready to go. On family night it keeps things age-appropriate and checks the content properly, because a cartoon isn’t automatically safe for the youngest one in the room.

In the demo I recorded, I set the scene for two of us, mood set to “bend my mind,” up to three hours. It came back with The Expanse, and instead of a lonely star rating it told me why: it shares a tone with The Blacklist, which I rated 8. Then, the part I love most, it gave me one honest reason I might pass. I trust a suggestion far more when it’s willing to argue against itself.

None of that lives in a rating prediction. All of it lives in the workflow around the decision. The Oracle was a distraction. The Concierge is the product.

WEB VERSION:

MOBILE VERSION:

The seam I still can’t stitch

Let me be honest about where this breaks.

Getting reliable streaming availability, country by country, is the one piece I still haven’t cracked. What’s actually on Netflix, or a local service, where I live and where I travel, changes constantly, and I haven’t found a source I trust. I tried a few tools, WatchMode among them, but none of them quite worked for the countries I care about.

So this is a genuine ask. If you know where to get dependable, country-level streaming catalogues, tell me. That single missing piece is what stands between “clever” and “I’d never watch anything without it.”

Build your own

If you want to try this, here is the shape of it. You don’t need an API bill or a computer science degree. You need your own ratings and a few careful hours.

  1. Export your history. Your IMDb ratings, or Letterboxd, or wherever you’ve quietly been keeping score. This is the fuel. Nine hundred ratings is plenty. A couple of hundred will do.
  2. Get the raw catalogue. IMDb publishes its non-commercial datasets for free. They are large but plain text, and a capable agentic model will unpack and join them for you without you touching a spreadsheet.
  3. Let the AI build the profile, then argue with it. Have it find the patterns in your own scores, then tell it where it’s wrong. Mine learned that a 6 from me means “barely watchable,” not “above average,” which changes everything downstream.
  4. Test before you trust. Hide some of your ratings, make the model predict them, and only then look. If a plain public average does nearly as well as your fancy model, believe the number. Then stop polishing prediction and start solving the actual decision: no repeats, who’s watching, what’s available where you are.
  5. Keep it private. IMDb’s free data is for personal use, so this powers your own tool, not a public product.

That last point isn’t something I can release to the world, and that’s fine by me. The how-to is the part worth sharing. The rest only works because it’s built around one specific household: mine.

Why this matters well beyond movie night

I spend my working life around AI in a very human industry, and this little experiment taught me something I keep seeing play out at far higher stakes.

We keep asking whether the model is smart enough. Usually that’s the wrong question. The model is rarely the moat. The moat is the boring scaffolding around it: knowing what’s already been done, who it’s for, what’s actually possible right now, and having the honesty to check whether the clever bit earned its keep.

My taste model didn’t earn its keep. The Concierge around it did. If I’d only celebrated the 12.8 million records and never run the test, I’d have shipped a good-looking lie to myself.

So the catalogue was never the point. The Concierge was.

Here’s what I keep coming back to. If an AI could pick your films perfectly, every single night, would you actually let it? I’m not sure I would. The best film I watched this year is one no model would ever have put in front of me. What was yours?


Discover more from Hotelemarketer by Jitendra Jain (JJ)

Subscribe to get the latest posts sent to your email.

0 comments on “Does AI Actually Know Your Taste? I Scanned 12.8 Million Movies to Find Out

Leave a Reply

Discover more from Hotelemarketer by Jitendra Jain (JJ)

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Hotelemarketer by Jitendra Jain (JJ)

Subscribe now to keep reading and get access to the full archive.

Continue reading