Film Diary
An Automated Movie Tracking & Analytics Platform
Case Study
Letterboxd rules. Full stop. But I like having a personal centralized system for all my stuff. I don't like being on a million platforms. And if the site ever disappeared, so would my entire film history since 2019. All of this is why I set out to create my own tracking system while still using Letterboxd.
This sounded easy at first. After all, Letterboxd already has all this information, so they probably have an open REST API, right? Nope. They're notoriously lock-and-key with their API access. But the site does mention that every user has an RSS feed. So I used Film Diary as a technical challenge:
Build the exact system I wanted with or without Letterboxd's cooperation.
The result is a fully automated movie tracking platform that has logged over 630 films since January 5, 2019. It captures watch events from my Letterboxd RSS feed, enriches each entry with comprehensive metadata from TMDB (40+ fields per film), and visualizes everything through beautiful interactive statistics, charts, and analytics.
What started as a workaround for a closed API has evolved into a sophisticated data pipeline featuring incremental updates, advanced caching, and a complete analytics dashboard. It's a love letter to film culture through the lens of data science — celebrating nearly a decade of cinema obsession with zero manual data entry.
Timeline and Development Process
Began Using Letterboxd
January 5, 2019
Initial System Build
May 2025
TMDB Profile Pages Integration
June 2025
Testing and Upgrades
Ongoing
Goals
- Other than logging a film in Letterboxd, no manual data entry required.
- Enrich watch logs with comprehensive TMDB metadata and cross-system integration.
- Create beautiful, interactive statistics and visualizations for exploring film history.
- Build a scalable caching system for optimal performance.
- Create a visually appealing and interactive personal film diary to use regularly.
What I Learned
Building Film Diary taught me what it really means to design a living system. I started with one simple idea—log my movies automatically—and quickly realized I was designing an entire ecosystem that needed to think for itself. Every decision rippled outward: one tweak to caching changed load times across the site, one small fix to the TMDB parser reshaped the accuracy of every statistic.
I didn’t hand-write the code that powers it all; I directed it. My Cursor agents built the foundation, but I was the one defining the blueprint—what belonged where, how it should behave, and what “done” actually meant. When things broke (and they did, often), I wasn’t debugging syntax so much as debugging logic: why a cache missed, why an update loop stalled, why a film didn’t resolve to the right metadata. Each iteration taught me more about how data wants to move through a system, and how to make those movements efficient without losing flexibility.
Working with agentic AI turned out to be a lot like working with actors. As a former film director and producer, I’ve been in countless moments on set where a line or scene just doesn’t land—no matter how many takes you shoot. A good actor will keep pushing, convinced they can nail it on the next try, because that’s their job: to make it work. But as the director, you have to recognize when it’s not about the performance—it’s about the setup. You pivot, you rewrite, you redirect. AI agents are exactly the same. They’ll try something forever if you let them, repeating the same logic loops, insisting they can fix it with one more run. The key is knowing when to step in, stop the take, and say, “We’re changing direction.” That’s what separates meaningful, human-guided work from algorithmic slop—and it’s one of the most important lessons I learned.
By the time it stabilized, I understood how to architect pipelines that could scale, cache systems that could think ahead, and analytics that could tell stories instead of just showing numbers. More importantly, I learned that technical direction is still a form of authorship. Every line of code might have been written by an agent, but the vision, structure, and flow of Film Diary came entirely from me.
All-Time Statistics Comprehensive analytics across 7 years of film watching - and counting.
What I Directed
Automated Data Collection Pipeline
The system monitors my Letterboxd RSS feed hourly via cron job, automatically capturing every film I watch. This means when I log a movie in my Letterboxd app as it starts, it appears on my site within the hour — fully enriched with TMDB's comprehensive movie database (cast, crew, genres, ratings, production details, and 40+ metadata fields) plus Movie Lists & Targets tracking data. What starts as a simple RSS entry becomes a rich, interconnected dataset across multiple systems.
Incremental Update System
Instead of processing the entire dataset daily, I directed my AI agents to assemble an intelligent incremental update pipeline that processes only new films. It sorts entries by watch date (newest first) and stops at the first existing entry — reducing daily updates from hours to under a minute while maintaining complete data integrity.
Advanced Caching Architecture
To eliminate redundant API calls and maximize performance, I architected a three-tier caching system—movie, director, and actor layers—and had the agents wire it into the PHP layer so pages load instantly without extra HTTP requests.
Rich Analytics Dashboard
The system features both all-time and yearly statistics pages with animated genre bar charts, top directors and actors with cached profile photos, weekly viewing patterns visualized with Chart.js, and chronological film lists organized by month. Person cards link to detailed profile pages, and film cards connect to TMDB movie pages — all responsive and optimized for dark mode.
52-week viewing pattern visualization with Chart.js and dark mode support.
Technologies Used
- Frontend: HTML5, CSS3 (modular design system), Vanilla JavaScript (ES6+)
- Backend: PHP 7.4+ with modular architecture and includes
- APIs: TMDB API for movie data, Letterboxd RSS for watch events
- Data Storage: JSON file-based system with yearly data files
- Visualization: Chart.js for interactive charts and graphs
- Automation: Cron jobs for daily RSS sync and TMDB enrichment
- Caching: Custom PHP/JavaScript caching utilities with class-based architecture
My Role
Responsibilities
- Generative AI director and systems architect overseeing the entire platform.
- Architected the complete data pipeline and directed AI agents to implement it from RSS ingest through interactive visuals.
- Directed the modular caching system and enforced the PHP injection approach that keeps performance tight.
- Orchestrated custom JavaScript utilities for statistics rendering and Chart.js visualizations.
- Defined the incremental update strategies and guided agents through integrating them with the sync cron.
- Continue steering the platform’s evolution—setting priorities, reviewing agent output, and safeguarding quality.
Person Cards Directors and actors displayed with cached profile photos and filmographies.
Modular Architecture Clean, maintainable code with reusable functions and class-based utilities.
Custom Features and Innovations
Flexible Name Matching
To handle directors and actors with name variations, I directed Cursor to implement a three-tier matching system—exact match, case-insensitive match, and partial match—that resolves differences automatically.
Smart Deduplication
The system tracks films by title + year keys and maintains separate watch dates for rewatches. This allows accurate counting of viewing events while preventing duplicate entries, handling edge cases like watching the same film multiple times in the same year or across different years.
Real-World Impact
Film Diary has been running for several months, but is able to track every film I've watched since 2019, tracking over 630 films with zero manual data entry. Hourly updates complete in under one minute, and page loads average under 2 seconds. It's become an essential part of my film watching ritual — every viewing automatically becomes part of a larger story, revealing patterns in habits, favorite directors, genre preferences, and viewing frequency over time.
Yearly Statistics Detailed analytics for each year including most watched genres, directors, actors, and more.
Challenges
- API Rate Limiting: TMDB's rate limits required implementing intelligent throttling (4 requests/second) and batch processing with delays.
- Name Deduplication: Directors and actors appear with different name formats across datasets, requiring flexible three-tier matching algorithms.
- Dynamic Theming: Creating Chart.js visualizations that adapt to theme changes while maintaining readability across both modes.
- Long-Term Scalability: Building an architecture that handles unlimited growth without performance degradation as the dataset expands year after year.
Next Steps
- Implement advanced filtering by genre, director, actor, decade, and runtime.
- Build a recommendation engine based on viewing history and preferences.
- Integrate a watchlist system.
- Migrate from JSON files to a database (MySQL/PostgreSQL) for improved query performance.
- Finish and launch the achievement system for viewing milestones and streaks.