We're looking for a Principal Data Scientist to serve as the technical lead across an entire agentic AI program. This role is ideal for a deeply experienced data science and AI leader who can set direction across many workstreams at once: owning the technical vision, guiding large teams through complex design and delivery challenges, and keeping the full AI/ML program coherent, high quality, and moving at pace.
In this role, you will be the senior technical authority over every agentic AI workstream in the program, spanning a large delivery organization. You'll set architecture and evaluation standards, manage technical delivery across streams, and act as the trusted voice on technical decisions with client leadership. You'll partner closely with executives to define the AI roadmap and make high-stakes decisions that determine how agentic AI scales across the organization.
Why This Role Matters
At Robots & Pencils, we design AI systems for a human world. Our name says it all. Robots and pencils means engineering paired with creativity, because every agent we ship has to work for real people in real workflows. That balance is baked into how we operate.
Every role here contributes directly to that mission. Here, you shape how AI systems integrate into enterprise operations, how teams move at real velocity, and how products create measurable impact for clients and the people they serve. We ship production-ready AI in 30 to 45 days. That pace demands people who take ownership, lead with craft, and care deeply about what they put their name on.
What You'll Do
Craft & Delivery
- Provide technical oversight of all agentic AI workstreams, from problem framing and design through evaluation, deployment, and ongoing operation
- Define the technical strategy and reference architecture for agentic systems, including multi-agent orchestration, tool and function calling, RAG patterns, vector databases, embeddings, and streaming responses
- Manage technical delivery across streams, identifying cross-team dependencies, resolving technical blockers, and keeping quality and velocity high as the program scales
- Set the standard for model development, experimentation, and the path from research to production across the program
- Establish rigorous evaluation frameworks for LLM and agent performance, including quality, safety, hallucination, cost, and latency, so progress is measured and not assumed
- Guide the design of scalable ML platforms, pipelines, and workflow orchestration that support event-driven, asynchronous operations at scale
- Ensure AI reliability, security, and scalability across deployed systems, including observability, monitoring, and debugging in production
- Bring an AI-forward mindset to your daily work, using tools like Claude, Cursor, and other modern AI assistants to ship higher-quality work at pace
Collaboration & Communication
- Serve as the technical face of the AI/ML program to senior client stakeholders, building trust through clear, honest communication on progress, risk, and tradeoffs
- Co-define the AI roadmap with executive leadership, operating as a peer in strategic technical conversations
- Translate complex technical concepts for executive, engineering, and business audiences, turning depth into decisions others can act on
- Align data science, engineering, product, and business teams so every workstream ladders up to shared priorities and measurable outcomes
Leadership & Influence
- Lead and influence a large, multi-team delivery organization, setting expectations for technical excellence across every stream
- Define and champion data science and AI engineering standards that shape how the program builds, evaluates, and operates AI systems
- Review and elevate the quality of work across teams, giving direct feedback that makes the work better
- Mentor technical leads and senior practitioners, developing the next generation of technical leaders
- Own high-stakes technical decisions with program-wide weight, balancing innovation, risk, and delivery speed
- Drive technical vision, defining not just what gets built, but how agentic AI practice evolves across the program over time
What You'll Bring
- 12+ years of experience in data science and AI/ML, with a record of taking AI systems from research to production at enterprise scale
- Proven track record leading technical delivery across multiple concurrent workstreams or teams, with influence that extends well beyond the immediate team
- Experience leading, or technically overseeing, large multi-team programs, including managing technical risk, dependencies, and delivery quality
- Exceptional stakeholder management and communication skills, including credibility with senior executives and client leadership
- Hands-on experience with LLM and agentic systems, including prompt engineering, function/tool calling, multi-agent orchestration, RAG architectures, vector databases, embeddings, and streaming LLM responses
- Deep expertise in model evaluation, experimentation design, and applied statistics, including evaluation approaches specific to generative and agentic systems
- Strong proficiency in Python and the modern data science and ML stack
- Expertise in MLOps and AI infrastructure, including model versioning, monitoring, deployment automation, and reproducibility
- Strong software engineering fundamentals, including system design, API design, code quality, and strong unit testing practices
- Experience with distributed systems, event-driven architectures, and workflow orchestration tools
- In-depth experience with AWS, especially the AWS GenAI offering (e.g., Amazon Bedrock); working knowledge of other cloud platforms
- Familiarity with both SQL and NoSQL databases, including scalable design patterns
- Working knowledge of AI governance, responsible AI, and compliance considerations in production environments
Helpful Extras and Unique Skills
- Experience supporting AI programs in regulated industries such as life sciences, healthcare, or financial services
- Experience in a consulting or client-facing technical leadership role
- Background in fine-tuning, evaluation automation, or agent safety and guardrails
You'll Do Well Here if You Are
- A doer. You see something broken and fix it. You'd rather move on clarity than wait for certainty.
- A fast learner who knows you don't know everything. The AI landscape changes weekly. You're senior enough to know better and curious enough to keep learning anyway.
- Direct in a way that makes the work better. You give honest feedback. You'd rather have the hard conversation than blow smoke.
- Obsessed with craft. You know genius is in the details. You ship exceptional, not perfect, and you don't put your name on work you wouldn't stand behind.
- Built for ownership. You honor commitments, admit mistakes fast, and back your teammates when a decision costs something. No handoffs, no finger-pointing.
- All in. You treat clients' businesses like your own. You take the work seriously without taking yourself seriously.
- Resourceful when the budget, timeline, or team is tight. Constraints don't slow you down. They sharpen you.
- Glad to be in the room with people who care as much as you do. Our teams average fifteen-plus years of experience. We hire people who push each other to do better work.