Movie Lists & Targets
Creating a Functional Film Discovery System
Case Study
I started watching movies again in 2025 and was having trouble finding good ones. Most modern tools push you toward either their own content or only the latest releases. I’m eager to discover cool gems from the 1980s, 1990s, and 2000s that I might have missed, including my favorite filmmaker's own favorite movies. I also want to explore esoteric "best-of" lists that seem tailored to very specific audiences.
So I created Movie Lists & Targets - a comprehensive film curation and analysis system powered by 600+ curated lists from critics, directors, festivals, and institutions around the world. It's really two interconnected tools: Movie Lists lets me browse any of these lists with advanced filtering and cross-referenced watch history, while Targets uses statistical analysis to identify consensus picks - the films that appear on the most lists, revealing what the film world collectively considers essential cinema.
I use these tools frequently to discover what to watch next. It's genuinely improved my film-watching life, and that was always the goal.
Timeline and Development Process
Concept
&
First Lists Import
Early 2025
Database Consolidation
Spring 2025
Advanced Filtering System
&
Targets Statistical Analysis
Summer 2025
Regular Daily Use
Present Day
Goals
- Direct the build of a comprehensive tool that allows me to discover films across 600+ curated lists and view those lists with my watch history.
- Import and organize 600+ curated film lists from critics, directors, festivals, and institutions.
- Create a single consolidated database for efficient data management.
- Implement an advanced filtering system with exclusion-based logic for precise film discovery.
- Build statistical analysis to identify consensus picks across all lists (Targets).
- Cross-reference with my Film Diary (2019-present) to show what I've already seen.
- Create a tool for regular personal use that genuinely improves my life.
What I Learned
Movie Lists & Targets taught me two things that completely changed how I think about building large systems.
First, I learned that once you truly understand how your own pipeline works, scale stops being intimidating. What started as a few lists from Letterboxd turned into a 600-list ecosystem with four custom Python tools handling scraping, consolidation, updates, and incremental rebuilds. And yet, once the system stabilized, I realized how simple it actually was to maintain. Now, if I want to pull a hundred new lists, process them, and update everything, I just run my tools, and within minutes the entire platform refreshes itself—no manual cleanup, no broken dependencies, no guesswork. I learned that with the right architecture and tools, massive systems can stay flexible and approachable.
The second thing I learned was how often AI agents can miss the forest for the trees. When I first asked my Cursor agents to design the database, they built it the hard way—each list contained every movie and every piece of metadata, over and over again. It technically worked, but it was bloated and inefficient. I asked, “Shouldn’t we just have a single list of all the movies, and then store the ListIDs for each one instead of repeating their data?” The agent paused, processed, and then said what it says so often (a phrase I’ve tried multiple times to ban): You’re absolutely right!
That moment stuck with me. Because the truth is, even as a beginner programmer, I was the one seeing the bigger picture. I didn’t need to know every Python detail to architect something smarter—I just needed to understand what the system should do and keep pushing until it did. Sometimes when the AI says, “You’re absolutely right,” it’s not just frustratingly hollow reassurance—it’s confirmation that intuition and direction still matter more than syntax.
What I Created: Two Interconnected Systems
Movie Lists
Browse & Discover
Browse 600+ curated film collections with advanced pill-based filtering, watch history tracking, and expandable inline movie cards.
Key Features
- 600+ curated lists (13 categories)
- 50MB+ consolidated database
- Pill-based exclusion filtering
- Watch history checkmarks
- Inline expandable movie cards
- Pagination & lazy loading
Targets
Statistical Consensus
Pure consensus analysis showing which movies appear on the most curated lists, ranked by frequency with category filtering.
Key Features
- Frequency-based rankings
- Visual consensus badges
- Category-level filtering
- Publication/org filtering
- Top 10/25/50/100/250 views
- Cross-referenced watch history
The Data Pipeline: Four Python Tools
Finding and integrating 600+ curated film lists required building a complete data pipeline. I directed Cursor to create four specialized Python tools that work together to automate the entire process from discovery to deployment.
1. Letterboxd List Scraper
Letterboxd is the largest trove of curated film lists online, so I needed a dedicated scraper. I directed the build of an automated tool using Selenium and BeautifulSoup to extract movie lists directly from Letterboxd, handling the site's infinite scroll by scrolling until all movies are visible. The scraper pulls movie titles, years, and slugs from the DOM, then matches each film to the TMDB API to get standardized IDs, and finally exports everything into structured CSV files with proper metadata. With no official API, this tool was essential to reliably collect clean, usable data from Letterboxd.
2. Consolidate Movie Lists
Takes all CSV files from the movie-lists directory and consolidates them into the single 50MB+ movie-list-database.json file. For each unique movie across all lists, it makes a single TMDB API call to fetch complete metadata (cast, crew, genres, runtime, ratings, production details, images, keywords). Movies are stored as objects keyed by TMDB ID, and each movie tracks which lists contain it via comma-separated list IDs with position data like "01-001-5,02-015-42". This creates the unified database that powers both Movie Lists and Targets.
3. Add New Lists
Adds new lists to the existing database without rebuilding everything from scratch. This tool reads the current database, processes only the new CSV files specified, checks if movies already exist (no redundant API calls), fetches TMDB data only for new movies, and merges the results back into the database. This incremental approach is essential for maintaining the system - I can add 10 new lists without re-processing all 600+ existing ones.
4. Update Movie List
Updates existing lists when they change - when lists add new movies, remove old ones, or reorder their rankings. The tool compares the new CSV against the existing list in the database, identifies additions/removals/position changes, updates the database accordingly, and maintains data integrity across the entire system. This ensures Movie Lists and Targets always reflect the current state of each curated list.
Together, these four tools create a maintainable pipeline: scrape new lists, consolidate into database, incrementally add more lists, and update when lists change. What could have been hundreds of hours of manual data entry became an automated workflow that scales to thousands of lists.
Part 1: Movie Lists — Browse 600+ Curated Film Collections
Movie Lists is the browser - the tool that lets me explore any of the 600+ curated film collections I've imported. The heart of the system is a massive collection organized into 13 categories, from the authoritative (Sight & Sound's Top 250, AFI's 100 Greatest Films) to the playful (Nic Cage Freakouts, Movies Where the Dog Dies). I've got everything from Cannes winners to Roger Ebert's Great Movies to genre-specific deep dives like Essential Horror and Noir Classics.
What makes this work is the consolidated database approach. Instead of managing 600+ separate files, I directed the system to consolidate everything into a single movie-list-database.json that contains every movie and every list. Each movie entry includes complete TMDB metadata - cast, crew, production details, posters, backdrops - plus a comma-separated string showing which lists contain it. The entire database is 50MB+ but loads efficiently with smart pagination and caching.
Movie Lists Interface Browse any of 600+ curated collections with pill-based filtering and instant watch history cross-reference.
Advanced Filtering That Actually Works
The filtering system uses a pill-based UI with exclusion logic, which sounds counterintuitive but works brilliantly in practice. Instead of "show me only Action movies," it works by exclusion: click a genre pill and it turns red (excluded), removing movies where all genres are excluded. Click again to include it back. This creates incredibly flexible filtering that feels natural once you understand the logic.
I can filter by 19 TMDB genres with a master "Unselect All" pill for quick resets, runtime ranges from "70 mins or less" to "More than 150 mins," release decades from Pre-1940s through the 2020s, TMDB rating ranges, MPAA ratings, language preferences, and production countries via dropdowns.
Exclusion-Based Filtering
Pill-based UI with color-coded exclusion logic. White pills mean included, red pills mean excluded. Click to toggle, use master pills for quick resets. The system creates incredibly flexible discovery patterns that feel natural despite sophisticated underlying logic.
Watch History Integration
Every movie I've watched since 2019 is tracked with exact dates. When browsing any list, I see green checkmarks on posters showing how many times I've watched each film. Click a movie to see the formatted dates I watched it. This cross-referencing is invaluable - I can browse the Sight & Sound Top 250 and immediately see what I've already seen versus what's still on my list. I can then easily sort the list by watch history so the movies I've seen get pushed to the bottom.
Inline Movie Cards
Click any movie poster and an expandable card appears below the row with a backdrop hero image, complete metadata (title, year, runtime, rating, MPAA), cast and crew details, genres, production info, plot overview, watch history, and a list of all curated lists containing that movie. Each list is a clickable link that takes you directly to that list. Everything is interconnected.
Part 2: Targets — Statistical Consensus Analysis
Targets answers a simple question: Which movies appear on the most curated lists? It's a pure consensus algorithm - no editorial bias, just cold math revealing what the film world collectively considers essential. The engine counts every movie's appearances across all 600+ lists and ranks them by frequency.
Take The Godfather as an example. It appears on 68 different lists - AFI 100, Sight & Sound, Roger Ebert's Great Movies, 1001 Movies to See Before You Die, and dozens more. That count becomes its ranking metric. Movies are sorted by list count first, then by TMDB user rating as a tiebreaker. This ensures movies with identical list counts are further sorted by quality ratings.
Targets Statistical Analysis Movies ranked by list frequency with visual consensus badges showing appearance counts.
Advanced Filtering for Consensus Discovery
Targets includes all the standard filters plus two unique features that make consensus analysis incredibly powerful. Want to see consensus only from Canon Institutions (Sight & Sound, AFI, TSPDT) while excluding Fun Lists? The category pills let you include/exclude entire list categories. Click "Awards" to exclude all Oscar/Cannes/BAFTA lists, and the analysis recalculates showing only movies from the remaining categories.
I can also filter by publication or organization. Exclude "Letterboxd" to focus only on institutional and critical lists. Include only "Sight & Sound" and "BFI" for British film canon. The possibilities are endless, and each filter combination reveals different aspects of cinematic consensus.
Detailed List Information Each list has a detailed information panel with the list name, description, date and source information.
Visual Consensus Indicators
Every movie poster shows a badge with its list count - "47" means it appears on 47 lists. It's instant visual feedback for understanding consensus strength. A dropdown lets me view the top 10, 25, 50, 100, or 250 movies, perfect for quick "what should I watch tonight?" queries or deep research sessions.
Like Movie Lists, watched films show checkmarks. But in Targets, seeing which consensus picks I've already watched versus which remain unwatched is incredibly useful for queue building. And just like Movie Lists, click any movie to expand a detailed card showing exactly which lists contain it.
The Tech Stack
- Frontend: Vanilla JavaScript (2,147+ lines), no frameworks - complex state management without abstractions
- Backend: PHP endpoints for TMDB proxy and database access with rate limiting
- UI System: Responsive grid layouts (5 columns desktop, 4 mobile), pill-based filtering with color-coded exclusion logic
- Data Storage: Single 50MB+ consolidated JSON file (movie-list-database.json) with TMDB enrichment
- Watch History: 7 yearly JSON files (2019-2025) with dates and rewatch counts
- Icons: Material Icons for consistent UI actions, status indicators, and navigation
- Automation: Daily cron jobs for TMDB updates and cache maintenance
Movie Card Details Expandable cards show complete metadata and which curated lists each movie appears in.
Real-World Impact
Instead of scrolling through Netflix's algorithm-driven recommendations, I browse Cannes winners or my favorite filmmaker's favorite movies. The watch history cross-reference saves me time and prevents accidentally rewatching something I saw six years ago. Targets reveals genuinely great films through pure statistical consensus - no editorial bias, just collective wisdom from 600+ curated sources.
Technical Achievements
- Merged 600+ disparate film lists into single queryable database with consistent TMDB metadata
- Built four specialized Python tools to automate the entire data pipeline for build, deployment, and regular updates
- Sophisticated exclusion-based filtering logic creating incredibly flexible discovery patterns
- Statistical consensus analysis identifying patterns across disparate sources
- Complete integration with my personal watch history including dates and rewatch counts
- Real-time TMDB integration with caching, rate limiting, and fallback strategies
- Fully responsive interface with pagination, lazy loading, and smart caching for thousands of movies
Watch History Integration Green checkmarks show films already watched with exact dates.
Key Challenges Solved
- Database Size & Performance: 600+ lists with thousands of movies created massive datasets. Solution: consolidated single JSON file (one HTTP request instead of 600+), pagination (100 movies per page), lazy loading for visible posters only. Result: instant loads even for large lists.
- Filter Logic Complexity: Users think inclusion ("show Action movies") but system uses exclusion. Solution: smart pill UI where visual state matches mental model - white pills feel "included," red pills feel "excluded." Result: intuitive filtering despite complex underlying logic.
- List-to-Movie Relationships: Tracking relationships across 600+ sources without full database. Solution: comma-separated list IDs embedded in movie objects like
"lists": "01-001-5,02-015-42,08-012-141"with JavaScript split/filter operations. Result: efficient lookups without database complexity. - Watch History Integration: Loading 7 years of watch data (2000+ movies) on every page load. Solution: lazy load only when needed, cache in memory for session, show visual indicators without heavy processing, compute detailed stats only when movie card opens. Result: essential feature without performance penalty.
Next Steps
- Expand to 1000+ lists with festival lineups (Venice, Berlin, Toronto, Sundance) and decade-specific "best of" collections
- Add advanced sorting by decade dominance, genre concentration analysis, and cross-list discovery recommendations
- Implement social features: shareable list URLs, CSV/Letterboxd export, custom "best of" list generation
- Build personal queue management with watchlist, streaming availability integration, and availability notifications
- Create public API for third-party integrations and broader community use