Founding Marketer
Marketing & Communications
About Datalab
Datalab trains models that read documents reliably at scale. The world’s most important information is trapped in PDFs, scans, and files that can’t easily be parsed, and getting it out correctly matters. From frontier AI labs like Anthropic to Fortune 500s like Siemens, Datalab is where businesses turn when extraction has to be right.
We hit an 8-figure run rate with a team of 7. We have hundreds of customers across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools - chandra, surya, marker, and lift - have 70,000+ GitHub stars, millions of monthly downloads, and broad developer mindshare. We’re backed by founding members of OpenAI, FAIR, and Hugging Face.
Role Overview
Marketing at Datalab already has real infrastructure behind it. We run a combined launch and content calendar with a defined playbook - channels, owners, and cadence for every release - and we ship two to three launches a week against it. No one here works on marketing full-time, so nobody owns the larger question of how we show up when a developer, a buyer, or an AI agent goes looking for a document parser.
You’d be our first dedicated marketing hire. You’d inherit a working system rather than a blank page, and the job is to make it run reliably, then make it considerably more ambitious. The natural place to build from is our open source work, which reaches a wide audience and currently does less for us than it could.
This is a hands-on role. On a team of 7, you’ll write the blog post, draft the tweets, restructure the docs page, brief the agency if we hire one, and pull the numbers yourself. We’re looking for someone who has done this work somewhere it was done well, and who wants to keep doing it rather than direct someone else who does.
There is strong potential for this to grow into a leadership role as we hire more people into the marketing team.
An increasing share of our buyers never touch a search results page; they ask a model. We want someone who takes seriously how Datalab gets recommended by LLMs and agents - how our site, docs, and benchmarks are structured, where models pull their answers from, and how any of it gets measured. Nobody has fully figured this out yet, and we’d like you to be the person who works it out here.
If you want to inherit a brand book, a content team, and a defined category, this isn’t the right role. If you want to shape how a technical category gets talked about, with real customers and real distribution to work from, it’s a good one.
Day to day you will:
A typical week might look like: shipping the launch post for a new open source model, ghostwriting a thread for one of our researchers, editing a vertical explainer on tracked-changes extraction for legal teams, restructuring a docs page so a model can actually cite it, answering a sharp question in Discord, and digging into why organic signups dipped last week.
Own our model launches, especially open source. Every release - chandra, surya, marker, lift, and what comes next - should land. You’ll build the assets (launch posts, benchmark write-ups, demos, docs), coordinate the team’s posts, work HN and Reddit and the ML community, handle press where it’s worth handling, and run the retro afterward so the next one lands harder.
Take over and scale the content engine. Inherit the calendar, tighten the cadence, and own the strategy and much of the writing. Educational content for developers evaluating us, and vertical-specific content for the buyers who need to see their own documents in the example - legal, healthcare, financial services, insurance, logistics, government.
Ghostwrite for the team. Our engineers and researchers know things worth writing about and mostly won’t write it themselves. You’ll draft in their voice, in a way they’re glad to put their name on, and turn the team into a distribution channel.
Make Datalab legible to LLMs and agents. Structure the site, docs, and benchmark content so that models recommend us when someone asks how to parse a PDF. Own the technical work behind it, figure out what “ranking” even means in this context, and build a way to measure whether it’s working.
Treat the website as our front door. Most of our revenue starts with someone landing on datalab.to. Own the narrative, structure, and conversion path of the site and the top of our docs, working with engineering to ship changes.
Own the demand channels and the budget behind them. SEO, paid, landing pages, conversion. Propose and manage the spend, and own the numbers end to end — partnering with our Business Operations hire on funnel reporting and attribution so we know which channels actually produce revenue.
Arm the sales team. Build the case studies, benchmark comparisons, one-pagers, and decks our AE and solutions engineers need in live deals. Our strongest proof points are customers with hard extraction problems; turn those into assets that help close business and double as vertical content.
Show up in the community. Our Discord and GitHub are the largest audience we own. Set the tone there, spot what people keep asking for, and turn recurring questions into documentation and content. Represent us at conferences, meetups, and on podcasts — and build the speaking and events motion, including getting our researchers in front of the right rooms.
Track where AI is going and translate it into positioning. Agents, evals, context engineering, the shifting shape of the document AI market. Bring us a point of view on what it means for how we talk about ourselves, not just a summary of the news.
Feed the loop back to product and sales. What messaging converts, what objections keep surfacing, what the community keeps asking for — bring structured signal to the founder and the GTM team on a regular cadence.
What success looks like
By month 3
You’ve owned at least one model launch end to end, and it outperformed the last comparable release on a metric we agreed on in advance.
The launch and content calendar is shipping on schedule without slipping, with the team contributing under their own names.
Baselines are in place: organic traffic, signups by source, LLM-recommendation visibility, and a first read on what’s working.
By month 12
Inbound is meaningfully up and attributable, with at least one channel you built from nothing producing consistent pipeline.
The launch playbook is documented well enough that a new hire could run a release with it.
Datalab shows up when a model is asked how to parse documents at scale — and you can prove it moved because of work you did.
Ideal candidate
You’ve built this before at a company where it mattered. You write well and fast, and you’d rather ship a good post today than a perfect one next month. You’re comfortable being measured - you’d rather have a number attached to your work than not. You’re ambitious about where this goes and want more scope over time, but you’re not waiting for it before you start doing the work.
You’re technical enough to hold your own. Our audience is engineers, ML leads, and CTOs, and they can tell instantly when marketing doesn’t understand the product. You should be able to read a benchmark table and know what it means, run our models yourself, and write something a skeptical developer on Hacker News would find useful rather than annoying.
You’re collaborative and low-ego. You’ll be pulling engineers into launches, asking researchers for time, and putting your writing out under someone else’s name regularly. You make the call that’s right for the company, and you’re glad when the work lands even if the byline isn’t yours.
We’re eager to work with someone who has:
7+ years in marketing, with meaningful time owning content, launches, and demand generation for a technical or developer-facing product.
Run launches at a high level - coordinated messaging, assets, press, and community across a real release cycle - and can point to both the results and the parts you personally shipped.
Built a content program from strategy through published work, and written most of it yourself.
Owned SEO and paid channels directly, including budget, with numbers on what moved and what didn’t.
Operated in early-stage environments where the playbook didn’t exist and the team was small.
Can use AI tools like Claude and Cowork effectively to streamline your work and build content systems.
Bonus points if you:
Have marketed an open source project or built inside a developer community (GitHub, HN, Discord, ML Twitter).
Have worked on document AI, OCR, IDP, or data extraction - or adjacent infrastructure and developer tooling.
Have run experiments in LLM/AI-assistant discoverability and have opinions about what actually works.
Have ghostwritten for founders or technical leaders and have samples you can talk through.
Are comfortable enough with design to get a launch asset out the door yourself, and know when to bring in help.
Have marketed into regulated or public-sector buyers, where the proof burden is higher.
Have hired contractors, agencies, or freelancers and gotten good work out of them.
Want to build and lead a marketing team, and can tell us why you think you’re ready for that next.
Are active in the AI community and actually use open source models and tools.
Interview process
A 30-minute video call with the founder to evaluate fit.
A 1-hour conversation going deep on a launch or content program you’ve owned end to end - what you did, what worked, what you’d change.
A 2-hour onsite in NYC, working through three things: traffic and funnel data, a live editing and ghostwriting session using our actual drafts and voices, and a discussion of a short strategy brief.
Final conversations with the team.
We expect you to use AI tools throughout, including on the prepared brief - we use them constantly and want to see how you work, not watch you pretend you don’t.
At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.