Treat content as data
Product descriptions are parsed into explicit attributes rather than stored as undifferentiated marketing copy. That makes discovery, comparison and programmatic pages possible.
A solo-built discovery platform that turns fragmented Indian specialty-coffee catalogues into structured, searchable data.
Founder, product owner and sole builder · Updated Aug 2026
01 · Problem
India's specialty coffee is spread across dozens of independent roaster sites with no shared catalogue or consistent product data.
02 · System
Scrapers collect source catalogues, language models normalize inconsistent descriptions into a common schema, and automated QA and publishing loops keep the catalogue useful.
03 · Outcome
92 roasters, 1,400+ SKUs, 2,000+ monthly uniques and 1,000–3,000 weekly active users through organic discovery.
Coffee buyers often know a flavour, process or origin they want, but not which roaster currently carries it. Roaster sites describe similar attributes in different ways, making comparison difficult.
The product had to work as both a consumer discovery experience and a continuously maintained data system. Manual catalogue entry would have made the project impossible to operate alone.
Product descriptions are parsed into explicit attributes rather than stored as undifferentiated marketing copy. That makes discovery, comparison and programmatic pages possible.
Automation handles repeatable extraction and refresh work. Review effort is reserved for ambiguous records and schema failures instead of every item.
Pages are organized around questions coffee buyers already ask, so the same structured data serves the interface and the search surface.
Connected work is resolved from the shared content graph.