
About Pencil
At Pencil, we are driving innovation in advertising technology through our state-of-the-art SaaS product, which harnesses Generative AI to redefine content creation. Our mission is to make AI the default in advertising without replacing creative people. To achieve this, we need to make sure that our technology isn’t just in the hands of big brands—we need to help small businesses and creative individuals too.
We’re building the agentic OS for marketing. We aren't just integrating generative AI; we are building the machine that makes it professional, brand-safe, and scalable. We’re moving beyond simple text-in/text-out interfaces toward complex multi-agent architectures that can handle the nuanced demands of global brands and the agility of small businesses.
We are looking for an Agent Architect who thinks in systems, not just sentences. You will be responsible for the "brain" of Pencil—designing the agent logic, tool-calling structures, and evaluation loops that power our core creative engine and bespoke client solutions.
The Role: Architecting the Creative Brain
You won’t just be "prompting"; you will be engineering behavior. You will bridge the gap between creative intent and machine execution, ensuring that our agents are robust, predictable, and capable of high-fidelity output across text, image, and video. This is a client-facing technical role: you’ll spend as much time in rooms with brand and agency stakeholders as you will in the system. You need to be able to lead a discovery conversation with a non-technical client, decompose their problem into an agent architecture, and explain your design choices in plain language.
Your work will fall into two high-impact pillars:
- Core Systems: Designing and scaling the foundational agents that power the Pencil SaaS platform.
- Client Solutions: Architecting custom workflows for world-class brands that require specific "brand DNA" and complex creative logic.
Key Responsibilities
- Agent Architecture: Design and implement multi-agent workflows, including task decomposition, state management, and tool-use (RAG, API integration, etc.).
- Systematic Optimization: Move beyond "vibe-based" testing. Implement rigorous evaluation frameworks (e.g., using LLM-as-a-judge, promptfoo, or DSPy) to measure and improve agent performance at scale.
- Scalable Frameworks: Develop reusable agents and prompt libraries that allow the platform to serve 1,000+ brands with unique voices simultaneously.
- Cross-Functional Engineering: Partner with AI Engineering and Product teams to determine which behaviors should be handled via prompting, RAG, or fine-tuning.
- Client Architecture: Act as the technical lead for complex client deployments. Run structured discovery sessions with clients, decompose ambiguous business problems into buildable agent workflows, and present trade-offs (cost, latency, reliability) to non-technical stakeholders in plain terms.
Your Background
- 3+ Years of Direct GenAI Experience: Deep, intuitive, and technical understanding of LLMs (GPT-4, Claude, Gemini) and multimodal models (Stable Diffusion, Midjourney, Video Gen).
- Systems Thinking: You don’t just write a prompt; you think about the latent space, the context window, and how one agent's output becomes another’s input.
- Technical Literacy: Comfortable with Python, JSON structures, and API documentation. Experience with agent orchestration patterns (task decomposition, tool use, state management) matters more than any specific framework.
- Evaluation Obsession: You believe that if you can’t measure a prompt’s performance, you shouldn’t ship it. You are familiar with benchmarking and A/B testing AI outputs.
- The Creative/Technical Bridge: Ability to sit in a room with a Creative Director and translate creative feedback into concrete system changes. Comfortable being the technical face in client meetings.
- Client-Facing Problem Decomposition: Demonstrated experience leading discovery with customers or stakeholders and turning vague briefs into scoped technical solutions (solutions architecture, forward-deployed engineering, or technical consulting backgrounds fit well).
- Plain-Language Communication: Ability to explain agent failure modes, eval results, and architectural trade-offs to executives without jargon.
You'll Thrive Here If...
- You find "hallucinations" to be a logic puzzle to be solved, not just a bug.
- You are excited by the challenge of making an AI follow a 50-page brand book with 100% fidelity.
- You want to build the infrastructure that defines how the next generation of advertising is created.
- You’d rather sit with a confused client and untangle their problem than receive a perfect spec.
Key Performance Indicators & Success Measures
- System Reliability: Reducing the failure rate of complex agent workflows.
- Architectural Efficiency: Minimizing token usage and latency while increasing output quality.
- Brand Alignment: Automated evaluation scoring of agent outputs against specific brand guidelines.
- Adoption at Scale: Core agents adopted and retained across the Pencil user base.
- Client Deployment Success: Custom client workflows delivered on time and adopted by the client.
- Stakeholder Trust: Clients and internal partners return to you to scope strategic accounts.
Benefits
- 25 days PTO plus public holidays (flexible time off scheme)
- Health insurance / private medical cover
- Monthly stipend towards wellness, fitness, and learning & development
- Remote work (work from anywhere in Mexico)
- Enhanced parental leave policies (birth, adoption, or surrogacy)
- Flexible working hours
Timezone overlap
UTC-6–-3
Benefits
Open to
LATAM
Sign in to track applications and earn points.