
About Poolside
Poolside exists to build a world where AI will be the engine behind economically valuable work and scientific progress. We believe the fastest way to reach AGI lies in accelerating software development itself, by reshaping the developer experience with agentic systems, coding assistants, and the frontier models that power them. We deploy these systems directly into the development environments of security-conscious enterprises.
About Our Team
We were founded in the US and have our home there, but our team is distributed across Europe and North America. We get our fix of in-person collaboration in Paris each month for 3 days, with an open invitation to stay the whole week. For those based in PST, we understand this is a significant travel cadence; we are open to agree on a lower cadence and will discuss this in the interview process. We also do longer off-sites once a year.
Our team is a multidisciplinary blend of research, engineering, and business experts. What unites us is our deep care for what we build together. We're in a race that requires hard work, intellectual curiosity, and obsession; to balance this intensity, we've assembled a team of low ego and kind-hearted individuals who have built the special culture Poolside has.
About the Role
You'll be working on our data team focused on the quality of the datasets being delivered for training our models. This is a hands-on role where your #1 mission would be to improve the quality of our datasets across the entire training cycle (pre-training, mid-training, post-training, RL) by leveraging your previous experience, intuition and training experiments. This role particularly focuses on generating synthetic data at scale and determining the best strategies to leverage such data into training large models.
Staying in sync with the latest state-of-the-art research in synthetic data generation and LLM training is key to success in this role. You will constantly lead original research initiatives through short, time-bounded experiments while deploying highly technical engineering solutions into production.
Your Mission
To deliver large, high-quality, and diverse synthetic datasets mixing natural language and code modalities to train best-in-class Poolside coding agents.
Responsibilities
- Follow the latest research related to LLMs and synthetic data generation in particular. Be familiar with the most relevant open-source datasets and models.
- Design and implement complex pipelines that can generate large amounts of data while maintaining high diversity and optimizing the resources available.
- Collaborate closely with cross-functional teams to ensure experiments run efficiently and data generated optimizes compute and time resources.
- Continuously measure and refine the quality of datasets being generated while validating final data strategies through quantitative data ablation experiments.
Skills & Experience
- Strong machine learning and engineering background
- Experience with Large Language Models (LLM), including:
- Understanding of how LLMs learn
- Data ablations and scaling laws
- Post-training techniques
- Training reasoning and agentic models
- Experience with implementing cost-efficient, complex pipelines to generate synthetic datasets at scale
- Experience with evals tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc.)
- Experience building trillion-scale pretraining datasets, data curation, deduplication, tokenization, curriculum, etc.
- Excellent programming skills in Python
- Strong prompt engineering skills
- Experience working with large-scale GPU clusters and distributed data pipelines
- Strong obsession with data quality
- Research experience (nice to have): Author of scientific papers on applied deep learning, LLMs, or source code generation
Benefits
- Fully remote work & flexible hours
- 37 days/year of vacation & holidays
- Health insurance allowance for you & dependents
- 16 weeks of flexible, full-pay parental leave
- Company-provided equipment
- Well-being, always-be-learning & home office allowances
- Frequent team get-togethers
- Diverse & inclusive people-first culture
Timezone overlap
UTC-8–-7
Benefits
Health, Parental leave, Equipment, Wellness, Learning, Home office, PTO
Open to
US
Sign in to track applications and earn points.