Work
Product leadership, and products built solo with AI agents
Product leadership case studies from Affable and Bazaarvoice. Covers the specific situation, the actions I drove, and the measurable outcomes.
Product Leadership
Natural-language creator discovery across an index of six million creators.
Creator Discovery was Affable's flagship product. Customer research revealed that users struggled with conventional keyword search, specifically the friction of toggling 20+ filters per session and the inability to find 'creators like this one' within an index of six million.
As the first product hire, I built AI search from concept to platform, delivering semantic, visual, multilingual, and lookalike retrieval. Post-acquisition, I focused on optimizing performance by establishing a rigorous evaluation framework and observability program to ensure enterprise-grade search quality.
Affable's search became the primary creator-sourcing engine inside Bazaarvoice's platform. The quality program took demographic model accuracy from 72% to 93% and cut latency 68%.
During the early stages of the creator discovery launch, search quality was assessed anecdotally, relying on client feedback and demo success rather than rigorous data. With an index of over six million creators, the relevance of the top twenty results for any query went unmeasured. This was a deliberate trade-off for a seed-stage team with limited product resources.
After the acquisition, Creator Search was the primary creator-sourcing engine inside an enterprise platform, which meant its quality needed to be measurable: a golden set of queries with agreed-correct results, an accuracy metric the team could interrogate, and an observability program to catch regressions before clients did.
The first measurement put demographic model accuracy at 72%. That baseline turned an open question into a prioritized backlog, with every failure case in the golden set becoming an item of work. The program closed at 93% accuracy, with latency down 68%. The same period shipped a creator brand-safety score on the GARM framework, reading image, video, and caption to flag risk at the point of discovery.
A creator brand-safety score that ranks by harm severity across image, video, caption, and comments.
Brand-safety scores help brands identify risky creators during discovery. The original model used a proportional scoring system, where risk was determined by the percentage of flagged posts in a feed. This meant a creator with one post inciting violence across 200 posts was ranked as "safer" than a creator with ten posts about alcohol. This ranking was driven by volume rather than the severity of the violation.
I replaced the proportional model with a new system where harm is weighted by severity, analyzed across all modalities, and aligned with each brand's specific risk policy.
The score now prioritizes harm: a single instance of incitement triggers a high-risk rating, while multiple instances of low-severity content (like alcohol) result in a low-risk rating. Severity, recency, and frequency now drive the score by analyzing images, videos, captions, and comments. The system also includes on-demand deep scans and human reviews for contested cases.
A proportional score treats every flagged post as equal and rewards a large, mostly clean feed. The question a brand is actually asking is closer to a gate: is there anything here that disqualifies this creator. One serious violation can be disqualifying on its own, however small its share of the feed, so the model needed severity to override volume rather than average against it.
The rebuild scored GARM harm categories weighted by severity, recency, and frequency, reading image, video, caption, and comments so that risk carried in a video or a comment thread was not missed by a caption-only pass. Brands set their own tolerances against the GARM categories, so the same creator can read safe for one brand and risky for another. An on-demand deep scan let a reviewer pull a closer read on a specific creator, and human review handled the contested cases the model flagged but could not settle.
Three tools, one unified platform: sampling, ambassadors, and paid creators.
Enterprise clients bought sampling, ambassadors, and paid creators as three disconnected tools, which meant three sales motions, three roadmaps, and no unified value story.
Authored the investment case that reset portfolio priorities and architected the unified offering: a 9.5M+ member community, ambassadors, and paid creators sold and operated as one managed product. Aligned Sales, Finance, and four engineering teams on a unified product definition.
Enterprise clients now buy content sourcing by community mix in a single managed product. It anchored a $16M YoY growth category.
The three tools operated on separate roadmaps with independent owners and conflicting client commitments. The investment case argued two things: that the portfolio's growth targets did not work with the offerings sold separately, and that fixing this meant stopping some funded work to make room. The second argument consumed most of the alignment effort, because a priority reset means walking back commitments already made by other owners, which had to be negotiated stakeholder by stakeholder.
We shipped a unified product that integrated our 9.5M-member sampling community with ambassador and paid creator workflows. Success relied on maintaining a consistent product definition across Sales, Finance, and four engineering teams. This required rigorous coordination to ensure that how we sold, priced, and built the product remained aligned throughout every planning cycle.
Bazaarvoice's first creator payments platform: strategy, build-versus-buy call, and the build itself.
Deals were stalling over a critical objection: the platform lacked the ability to pay creators. The company had not moved money before.
I defined the payments strategy and roadmap, creating a build-versus-buy decision framework. After making a strong case internally, I led the end-to-end launch of Bazaarvoice's first creator payments platform (admin-led managed transfers, full transaction history) across Legal, Technical Services, FinOps, and Finance.
It addressed 28% of previously lost Affable deals and unblocked 24 stalled commercial conversations. The platform now serves as the core payment infrastructure for Integrated Sampling.
The build-versus-buy framework assessed long-term strategy rather than immediate needs. Payments wasn't just a one-off checkbox; within 18 months, we would need to handle sampling payouts, ambassador incentives, and affiliate commissions and a new vendor integration would have to be redone for each of them.
There was reasonable pushback. Finance preferred a specific vendor, Legal preferred the smallest possible compliance surface, and a credible vendor already existed for this exact use case, so taking on money-movement risk voluntarily required strong justification. The framework resolved it on a narrow point: the vendor route solved the immediate deal and left the roadmap use cases of payouts, incentives, and commissions, each requiring its own re-integration.
The build scope covered admin-led managed transfers, full transaction history, and a new compliance path defined with Legal and FinOps. After launch it addressed 28% of previously lost Affable deals and reopened 24 stalled commercial conversations, and the platform now serves as the payment rail under Integrated Sampling.
Transitioned a specialized product team to AI-native workflows.
Our roadmap shifted toward AI features faster than our team's operational model could support. Because this transition wasn't a planned initiative, I had to advocate for and implement changes incrementally.
I transformed the team's operating model by integrating AI prototyping into the scoping phase and establishing a formal evaluation discipline. I also set the roadmap for agentic workflows, backed an agent prototype ahead of engineering capacity, secured necessary LLM tooling, and coached PMs through their first AI-driven feature launches.
The team now standardizes AI feature development with built-in evaluation processes. For Intelligent Product Tagging, we improved model accuracy 6x, while maintaining the latency, within one year successfully hitting all phased OKRs.
This transition was executed as a series of deliberate, incremental adjustments. First, I secured essential LLM tooling. By embedding AI prototyping into our scoping process, feasibility questions were answered before commitment. An agent prototype was built ahead of engineering capacity, accepted as an unresourced bet so the team would have a working reference point when capacity opened up.
The effectiveness of this model is best demonstrated by our work on Intelligent Product Tagging. Previously, accuracy stalled at 10% with no framework to diagnose failures. Under the new operating model, we built a 'golden set' of queries, analyzed failure cases, and treated every prompt revision as a testable hypothesis. This methodology drove accuracy to 60% within a year across 25 enterprise pilot clients, establishing a new default standard for the team's AI development.
Built and launched a transactional video-on-demand platform in weeks to replace cinema revenue during the pandemic.
March 2020: Cinemas shut down and the company's core business stalled. We needed a functional transactional streaming service launched within weeks.
Led the end-to-end launch of transactional VoD across web, mobile, and TV. Owned discovery, checkout, and the cross-device experience, frequently pivoting the roadmap to keep pace with rapid market shifts.
The platform became the company's primary revenue driver while theatres remained closed, accounting for ~80% of total revenue with 30% MoM growth. Simultaneously, we reduced support voice volume by 30% through automated Help and chat integration.
When cinemas closed in March 2020, the company's primary revenue stream vanished overnight. We had to build and launch a transactional streaming product on web, mobile, and TV in a matter of weeks. The typical quarterly planning cycle was condensed into days, requiring us to re-evaluate our product priorities weekly as the market evolved. Navigating this volatility meant continuously refining our discovery and checkout experiences to maintain momentum.
The platform ultimately generated approximately 80% of the company's revenue, sustaining a 30% month-over-month growth rate while theatres remained closed. Furthermore, I optimized the customer experience by reworking the Help section and introducing contextual chat; this reduced support voice volume by 30%, effectively scaling our service capacity without additional headcount.
Built with AI
I build these products solo: I design the strategy, write the specs, and direct AI agents through the engineering. Each project includes a decision log and a comprehensive test plan. I add new projects as they ship.
A plain-language dictionary for AI jargon: a web app and Chrome extension that underlines terms as you read and explains any sentence you highlight, grounded in the glossary.
Try it live at aidecoder.app → GitHub →
Notes on shipping a retrieval stack
Most teams reach for a bigger model when the answers go wrong. In our case the answers were wrong because the context was wrong, and no amount of model changed that.
We reduced hallucination by tightening the retriever's top-k and reranking before the context is assembled. That single change did more than a month of prompt work.
Two columns
The embedding index was rebuilt offline every night, which meant a stale card could sit in production for a day. Fine-tuning was never on the table; the failure mode was retrieval, not generation.1
Chunking is where the rest of the tuning time went. A chunk that spans two ideas retrieves for both and answers neither, e.g. a paragraph that opens on latency and closes on cost.
A code block
The guard below runs before the model call, so a refusal costs nothing:
if (sentence.length > MAX_SENTENCE) return send(res, 400, { error: 'sentence_too_long' });
An agent that plans its own retrieval loop is a different product. We kept the pipeline fixed so that two runs stay comparable.
A table
| Change | Effect |
|---|---|
| Reranking on | Precision up, latency up by 180ms |
| Top-k 8 → 5 | Fewer distractors in the prompt window[2] |
1. Measured on 39 rows. 2. The reader never sees the retrieved set.
A dark section
Quantization let the model run on one card, and the quality loss was smaller than the latency win. We shipped it the same week.
This article is awkward on purpose: two columns, a code block, a table, footnote markers, and a word split across a line. A selection that picks any of that up is repaired to one clean sentence before it is decoded. A selection inside the code block offers no Decode button.
Extraction runs live in your browser. The explanations are captured from the model at aidecoder.app, where the extension decodes any sentence you highlight.
The explain feature runs on grounded retrieval: it pulls the few glossary cards a sentence depends on, and a small model restates the sentence in plain English from only those cards. When the glossary doesn't cover a term, the answer says so.
A gold set of hand-answered sentences grades every prompt change, and one invented answer fails the run. The scores drive the decisions that matter: which model ships, when to explain a sentence or decline it, and how grounding is enforced.
Designed, spec'd, and QA'd by me; engineered by AI agents against my decision log and test plan.
Local-first personal finance that ingests Gmail statements and categorizes spending with an LLM, entirely on-device.
LLM categorization uses a confidence gate: the system flags low-certainty rows for human review. A golden-set evaluation suite validates every prompt change prior to deployment.
Corrections are logged append-only and fed back as few-shot examples. The suite runs 100+ automated tests. The same evals discipline from the Bazaarvoice work, applied solo.
Designed, spec'd, and QA'd by me; engineered by AI agents against my decision log and test plan.
Deleted the local database out of an old demo-phase habit; it held the real encrypted Gmail token. No code failed. The gap was in process: demo-phase habits had not been retired when the build started touching real state. The rule changed the same day, and the incident is recorded in the decision log under its own number.
Amazon price intelligence in the browser: daily checks, below-target alerts, trend forecasts.
A dual-mode parser (DOM, plus HTML-regex for the service worker) reads product pages. Linear-regression forecasting runs over price history, backed by a full logic-test suite.
The edge cases came from testing against real product pages; that is how the phantom-price bug below was caught.
Designed, spec'd, and QA'd by me; engineered by AI agents against my decision log and test plan.
An unavailable product was recording a related-product carousel's price as its own, and one false below-target alert fired before it was caught. Diagnosed against the real Amazon page. The fix deleted the page-wide price fallback entirely: when the page carries no reliable price, recording nothing is the correct outcome.
— in Mumbai, India