
At Sword, we’re building AI to heal billions and unlock humanity’s full potential. In doing so, we’re pioneering AI Care, a fundamentally new approach to healthcare built for medical reasoning, safety, and real-time treatment, not generic technology applied after the fact. As both a clinical-centric frontier AI lab and an applied AI platform, Sword is reimagining how care is delivered at scale, removing traditional barriers like appointments, waiting rooms, and stigma so more people can access the care they need.
Since 2020, Sword has expanded across physical therapy, women’s health, cardiometabolic, and mental health, moving beyond traditional sessions to a fully AI-native, 24/7 care program. More than 700,000 members across three continents have completed over 10 million AI sessions, helping 1,000+ enterprise clients avoid more than $1 billion in unnecessary healthcare costs. Backed by 42 clinical studies, 44+ patents, and more than $500 million raised from leading investors including Khosla Ventures, General Catalyst, and Founders Fund, Sword is defining a new standard for healthcare.
AI Proficiency at Sword
AI fluency is a core expectation at Sword. Every candidate is assessed against our three-level framework:
- Explorer (Level 1): Uses AI daily to boost personal productivity
- Builder (Level 2): Creates workflows and tools that elevate the whole team
- Integrator (Level 3): Embeds AI into products and processes at scale
Every hire must demonstrate at least Level 1 fluency.
What You'll Be Doing
- Own ML projects end-to-end: Take problems from exploration through production deployment and iterate based on real user interactions.
- Build agentic LLM systems: Design multi-step workflows with tool use, retrieval, and orchestration, ensuring clinical-grade reliability.
- Drive evaluation as core engineering: Build eval datasets, offline and online harnesses, LLM-as-judge pipelines with human review, and regression tests to prevent quality drops.
- Improve model quality: Apply prompting, retrieval, distillation, or fine-tuning based on empirical evidence.
- Work across the full AI stack: Manage data preparation, model adaptation, serving, monitoring, and feedback loops in production.
- Partner cross-functionally: Translate clinical requirements into technical decisions alongside Product, Clinical, and Engineering teams.
- Engineering leadership: Review code, share technical knowledge, and mentor engineers earlier in their careers.
Requirements
- Proven experience shipping production ML systems that users depend on.
- Hands-on production LLM experience: prompting, retrieval, tool calling, and agentic workflows.
- A rigorous approach to evaluation: experience building eval datasets and frameworks to distinguish real improvements from noise.
- Strong ML fundamentals with clear reasoning around architectural tradeoffs.
- Ability to navigate ambiguity and transform loosely defined problems into production systems.
- Solid software engineering skills, including production-quality code, familiarity with distributed systems, and pipeline debugging.
- Clear communication skills across technical and clinical stakeholders.
Bonus Points
- Experience with fine-tuning or preference optimization (RLHF, DPO, or similar).
- Background in healthcare AI or other high-stakes domains where model errors carry real cost.
- Track record of building agent frameworks or evaluation tooling from scratch.
- Contributions to open-source projects, technical writing, or active knowledge sharing.
Timezone overlap
UTC+0–+3
Benefits
Open to
Europe
Sign in to track applications and earn points.