We are a CS research lab led by Shreya Shankar, moving to Carnegie Mellon University in 2027. Our goal is to make unstructured data (e.g., free-form text, documents) as easy to process, explore, and understand as relational data. We take a user-centered and full-stack approach — we talk to people and build open-source data systems and user interfaces.
We have three projects going on:
- DocETL, an agentic map-reduce system for processing large unstructured datasets with LLMs (with our friends at the EPIC Data Lab at UC Berkeley).
- Quail, an open-source execution engine for DocETL and AI-SQL, with open-source models, that KV-cache-hitmaxxes and throughputmaxxes on GPUs.
- DocWriter, an IDE for AI-assisted writing (or is it human-assisted AI writing?).
We are hiring! We are looking for highly motivated and ambitious database, HCI, or AI researchers. PhD applicants: apply through the CMU CSD or HCII PhD applications. Postdocs, undergrads, and masters students: email shreyashankar@cmu.edu with why you want to work with us, your CV, and transcript. Don't write your email with AI, please!
Want to use our systems? We love to work directly with users. If any of the projects above are interesting to you as a user, please reach out to shreyashankar@cmu.edu (e.g., if you are a medical researcher and want to run LLMs on all your documents; if you are a journalist who wants to write with AI without ceding your agency). We run user studies year-round.
Blog
Coming soon.