Applied featureDirector✓ comp disclosed2 hours ago
Senior design leader for Admin UX and design system at legal AI platform; not an AI PM role.
- describes Harvey as 'frontier agentic AI' and 'AI-native' product
- mentions 'AI tooling for the design org' and deciding 'how Harvey's designers use AI'
- no concrete technical PM responsibility named: no mention of evals, model selection, latency budgets, data labelling, or model quality ownership
- role is design leadership, not AI product management; AI appears as context for design decisions, not as owned artefact
Applied featureLead✓ comp disclosed2 hours ago
Design leader for legal AI platform's admin UX and design system; no AI technical ownership.
- describes Harvey as 'frontier agentic AI' and 'AI-native' repeatedly
- mentions 'AI tooling for the design org' and deciding 'how Harvey's designers use AI'
- no concrete technical artefact owned: no mention of evals, model selection, latency budgets, data labelling, or failure modes
- role is design leadership for Admin UX and design systems, not AI product ownership
- AI appears as context and tooling surface, not as a technical responsibility
Applied featureLead14 hours ago
Audio Engineering Lead managing production team for ElevenLabs' managed dubbing/audiobook services.
- Role is audio engineering leadership for a managed services team, not PM
- Mentions 'custom voices' and 'cloning and generation quality' but no ownership of model selection, evals, or technical PM artefacts
- Focus is on hands-on audio production, DAW proficiency, and client management—not product management
Applied featureSenior IC1 day ago
Full-stack engineer prototyping new ChatGPT features; product-minded but not AI PM.
experimentation
- Build full stack product experiences that unlock value from new model capabilities
- Work on a new ChatGPT experience one week, explore a new model capability the next
- Use user feedback, product signals, and experimentation to determine which ideas should scale
- No mention of evals, model selection, latency budgets, data labelling, or any concrete technical AI artefact the PM owns
Applied featureLead1 day ago
Design leader for Scale's enterprise AI applications, focusing on human-agent interaction patterns.
human-in-the-loopagent behaviourhallucination ux
- Shape Scale's approach to human-agent interaction
- interface is where trust is built or lost
- interaction patterns for how enterprise users work alongside AI agents
- building AI-native workflows
- designing for agentic platforms
AI infraSenior IC1 day ago
Backend Product Engineer building orchestration and security infra for AI agents in enterprise cloud environments
- running agents in customer infrastructure
- integrating with source control, CI, and other developer tools
- Design reliable orchestration for long-running, parallel work
- Build security into execution workflows
- no mention of model selection, evals, fine-tuning, data labelling, or model quality ownership
Applied featureSenior IC✓ comp disclosed3 days ago
Security engineer building guardrails and red-team testing for Notion's AI agent and tool execution surfaces.
hallucination uxprompt engineeringai safety policyhuman-in-the-loopagent behaviour
- Build automated red-team and regression testing for risks such as prompt injection, indirect instruction following, data exfiltration, tool misuse
- Hands-on experience with prompt-injection testing, red teaming, model behavior evaluation, or abuse-resistant application design
- Define and build security architecture for product surfaces that operate across customer workspace content, including tool execution, content writes, retrieval, permission checks
AI infraStaff✓ comp disclosed3 days ago
Staff PM building data infrastructure and marketplace systems that power frontier AI model training and evaluation
eval designdata labellinghuman-in-the-loopml metricscost modellingexperimentation
- building the systems that directly shape AI model quality
- define how tasks are designed, how contributors interact with multi-turn chat interfaces, and how in-task quality is measured and enforced
- evaluation frameworks, and human-feedback systems that make models smarter, safer, and more capable
- owns the end-to-end tasking product for training the next generation of models
- set pay rates methodically by skill and geography, and design incentive structures that balance cost efficiency, data quality
Model / platformLead✓ comp disclosed3 days ago
Finance & Strategy lead owning model monetization pricing and enterprise product economics at Anthropic.
cost modelling
- owns model monetization: pricing, tiering, consumption mechanics across model families
- connects token economics, inference and compute costs to pricing decisions
- no mention of eval design, model selection, fine-tuning, or any technical ML artefact
- role is financial/business strategy, not product management of AI capabilities
Model / platformLead✓ comp disclosed3 days ago
Lead partnerships and ecosystem strategy for OpenAI's platform integrations with ChatGPT and Codex.
- leads integrations for ChatGPT and Codex
- work with product and engineering on connectors, agents, and third-party workflows
- no mention of model evaluation, fine-tuning, data labelling, latency budgets, or any concrete technical artefact the PM owns
Applied featureSenior IC3 days ago
Manage production teams and QC for ElevenLabs' managed dubbing/audiobook service in Arabic.
- ElevenLabs is an AI voice company, but this role is managing human production teams and quality control for dubbing/audiobooks
- mentions 'working with the Productions product and engineering team to optimize our tools and processes' but names no concrete AI artefact the PM owns
- no mention of model selection, evals, fine-tuning, latency budgets, or any technical AI responsibility
- role is fundamentally about onboarding linguists, QC, and editorial direction—not AI product decisions
AI-native 0→1Lead✓ comp disclosed4 days ago
Lead PM building Claude Science, an AI-native research workbench 0→1, owning evals and model behavior for scientific workflows.
eval designmodel selectionprompt engineeringagent behaviourhuman-in-the-loopml metricsexperimentation
- build and shape evals grounded in real scientific workflows
- get in the weeds of model transcripts and evals
- define target behaviors, build and shape evals grounded in real scientific workflows, surface failure modes from real usage, and feed them back into model development
- Have personally built evals or benchmarks for model capabilities, ideally agentic or scientific ones
- comfortable going deep on model behavior, prompting, and evaluation methodology
- coordinates multi-agent workflows with built-in review for citation and calculation errors
Applied featureLead4 days ago
PM shipping AI agents for enterprise financial services customers at Sierra platform
agent behaviourhuman-in-the-loopexperimentationml metrics
- build and ship AI agents that handle thousands of customer conversations a day
- work with technical counterparts to address technical challenges in business process
- strong technical fluency and ability to reason about AI systems, integrations, and failure modes
- experience with product development for AI agents a plus
- Develop and improve Sierra virtual agents to fit and anticipate customers' needs
Model / platformLead✓ comp disclosed6 days ago
PM owning Claude's multi-cloud safety, fraud, and compliance posture across AWS, GCP, Azure
ai safety policyhuman-in-the-loophallucination uxexperimentationml metrics
- Own the multi-cloud strategy and roadmap for our safeguards — data retention, automated review, human review, and enforcement
- Design product experiences and operations paths for managing eligibility criteria, verification paths, and the controls attached to differentiated capabilities
- Own fraud and abuse defenses for partner-sold offerings: bot and mass-registration abuse, pre- and post-sale risk controls
- Define the fraud signals we exchange with each cloud partner and turn them into product requirements
- Drive our transparency commitments to customers and partners on data access, review, and enforcement
- equally fluent in policy nuance and systems detail, and can explain the tradeoff between them
Applied featureLead7 days ago
Design leader managing 8+ designers on ChatGPT consumer growth surfaces; AI context but no AI PM depth.
- mentions 'AI-native consumer growth' and 'rethink familiar growth patterns and create experiences that feel truly native to AI'
- no concrete AI technical artefacts named: no mention of evals, model selection, latency budgets, data labelling, or model quality ownership
- role is design leadership for growth surfaces, not PM ownership of AI systems
Internal AI opsSenior IC✓ comp disclosed8 days ago
Growth Engineer building technical growth products and website tooling; AI is a surface-level ingredient, not core PM responsibility.
- Agents that pull from our own AI stack, Salesforce, and marketing tools to surface key intent signals
- leverage Decagon's own AI stack to create novel prospect experiences
- Experience with AI/LLM-powered tooling in a growth or marketing context (nice-to-have)
- No mention of model selection, evals, fine-tuning, data labelling, or any concrete AI artefact ownership
Applied featureLead✓ comp disclosed8 days ago
Design leader for Scale's enterprise AI applications, focusing on human-agent interaction patterns.
human-in-the-loopagent behaviour
- Shape Scale's approach to human-agent interaction
- designing for agentic platforms, and excited to shape the future of interaction patterns for AI and human-in-the-loop systems
- the interface is where trust is built or lost
- no mention of eval design, model selection, data labelling, latency budgets, or concrete technical artefacts the PM owns
AI-native 0→1Lead✓ comp disclosed8 days ago
Lead 0-to-1 product development for Claude deployments in nonprofits and public agencies; own strategy, prototyping, and access programs.
prompt engineeringexperimentationhuman-in-the-loop
- own two things: the 0-to-1 moonshot bets that could provide a step change in abundance, and the programs that get Claude into the hands of the organizations
- Build prototypes yourself to validate ideas before committing resources
- Lead 0-to-1 product development from research to internal prototypes to shipped products
- Define product strategy for experimental initiatives that push beyond our current offerings
- Design and run the programs through which beneficial organizations across all pillars get access to Claude
- Own the path from access to success. Find out what stops these organizations from deploying Claude and remove it
- Stay hands-on with emerging research and prototype with AI tools like Claude Code
ai safety policyLead✓ comp disclosed8 days ago
Lead PM for AI safety/safeguards systems, owning evals, detections, and risk mitigation across Claude products
eval designhallucination uxai safety policyml metricshuman-in-the-loop
- own the ideation, design, development and deployment of Safeguards systems
- develop detections, evals, interventions, and tools to measure and mitigate deployment and user risks
- Ability to write safety evals and communicate externally about safety
- Lead the development of metrics to understand the area, performance, blindspots
- designing and building state of the art safety systems
- designing and building metrics to evaluate risks, system performance, user impact
Applied featureSenior IC✓ comp disclosed8 days ago
Growth PM for Claude consumer product; standard acquisition/retention/monetization focus, no AI technical ownership.
experimentationml metrics
- Claude is described as 'one of the fastest-growing AI products' but PM role focuses on standard growth metrics: acquisition, activation, retention, monetization
- No mention of model quality, evals, fine-tuning, latency budgets, or any concrete AI technical artefact the PM owns
- Responsibilities are generic growth PM: 'Analyze product metrics', 'A/B testing', 'funnel optimization', 'user research'
- 'Balance rapid iteration with our commitment to AI safety and ethics' is stated but not operationalized into PM responsibilities
Applied featureSenior IC8 days ago
Senior IC building and deploying conversational AI agents for Fortune 500 customers, owning end-to-end agent design and production optimization.
agent behaviourprompt engineeringhuman-in-the-loopexperimentation
- Design, build, and optimize enterprise-grade AI agents in production
- going deep on each customer's workflows, pain points, and goals
- understand and shape AI agent designs and hold your own in a technical conversation with both engineers and executives
- Turn the frameworks and playbooks you develop into reusable assets
ai safety policyLead9 days ago
PM owning safety frameworks, evals, and metrics for multimodal model deployments at OpenAI
eval designai safety policyml metricshuman-in-the-loop
- Develop comprehensive frameworks for understanding and mitigating deployment safety risks, drawing on data analysis, expert consultation, and adversarial assessments
- Create scalable approaches for evaluating and improving safety outcomes
- Develop and continuously refine clear, actionable metrics that effectively capture safety performance and user experience at scale
- Establish repeatable processes to integrate cutting-edge AI safety research into OpenAI's models and product offerings
- Drive initiatives which ensure that OpenAI's audio, image, and video deployments are safe
Applied featureSenior IC✓ comp disclosed9 days ago
Product Designer shipping Claude UI/UX; aware of model capabilities but no AI PM technical ownership.
- says 'designing around capabilities that are emerging in real-time' and 'stay close to the models'
- mentions 'pay attention to where model capabilities are heading'
- no concrete technical responsibility named: no mention of evals, model selection, latency budgets, data labelling, or failure handling
- role is design-focused, not PM; no ownership of model quality outcomes or artefacts
AI infraStaff9 days ago
Solutions Architect role (not PM) advising on Gen AI/LLM deployment and fine-tuning for Databricks platform customers.
fine tuningagent behaviourmodel selectionprompt engineering
- Hands-on experience with LLMs and modern Deep Learning: Transformer architectures, PyTorch, and distributed/multi-GPU training
- Hands-on experience taking LLM post-training beyond fine-tuning: SFT plus reinforcement learning–based methods — reward modeling, RLHF/RLAIF, or policy optimization (e.g., PPO, GRPO)
- Work closely with specific accounts on our most challenging Gen AI use-cases around reinforcement learning and fine-tuning in particular
- However, this is a Solutions Architect/Field Engineering role, not a PM role—no ownership of product decisions, roadmap, or metrics named
AI infraStaff9 days ago
Solutions architect advising enterprise customers on production Gen AI deployments on Databricks platform
fine tuningrag vs finetuneagent behaviourmodel selectionprompt engineering
- hands-on experience taking LLM post-training beyond fine-tuning: SFT plus reinforcement learning–based methods — reward modeling, RLHF/RLAIF, or policy optimization
- work closely with specific accounts on our most challenging Gen AI use-cases around reinforcement learning and fine-tuning in particular
- advise customers on GenAI and Deep Learning architectures, translate requirements into technical solutions
Applied featureLead10 days ago
Lead PM building learning features in ChatGPT; strategy-focused, light on AI technical ownership
experimentationhuman-in-the-loop
- Partner with research teams to explore how AI can better advance learning outcomes
- build ChatGPT into a true learning agent
- Use both data and user feedback to validate impact
- ensure products are safe, effective, and aligned with real-world needs
Applied featureLead10 days ago
Design leader for Codex/ChatGPT B2B growth experiences; no AI technical ownership.
experimentationhallucination ux
- Translate complex and rapidly evolving AI capabilities into clear, approachable experiences
- Help define what AI-native growth experiences should look, feel, and behave like
- Are fluent in using AI as part of how you explore, prototype, communicate, and create
- No mention of model selection, evals, latency budgets, data labelling, or any concrete technical AI artefact the PM owns
Internal AI opsSenior IC✓ comp disclosed10 days ago
Data scientist driving analytics and experimentation across Anthropic's products and operations.
experimentationml metrics
- says 'AI/ML products, large language models' as a nice-to-have, never as core responsibility
- no mention of model evaluation, fine-tuning, evals, latency budgets, or data labelling
- role is general data science: 'define core metrics, build measurement frameworks, A/B testing'
- AI context is company mission, not the PM's technical ownership
Applied featureLead11 days ago
Lead PM building enterprise AI agents for government institutions; owns agent design, evaluation, and deployment.
eval designrag vs finetuneprompt engineeringagent behaviourhuman-in-the-loopexperimentation
- owns outcomes from concept through production for AI agents
- experience building or deploying AI/LLM products, especially conversational or long-horizon agents, and familiarity with evaluation, retrieval, prompt engineering, or agent tooling
- design proactive and long-horizon agents that keep constituents informed, collect missing information, and guide them through multi-step government processes
- turn patterns in constituent conversations into actionable insights
- coordinate across customer stakeholders, engineering, design, research, security, and operations
Internal AI opsSenior IC✓ comp disclosed11 days ago
Product Policy Manager owning safety risk assessment frameworks and evaluations for Anthropic product launches.
eval designhallucination uxhuman-in-the-loopai safety policyexperimentation
- Design and run bespoke evaluations for products that require tailored assessment
- Analyze the potential for misuse, unintended consequences, and harmful outputs of new model and product capabilities
- Develop and maintain risk assessment frameworks to identify and evaluate potential safety risks
- Conduct comprehensive product safety reviews, covering technical and non technical harms
- Leverage SME risk assessments to inform overall product safety recommendations
Applied featureDirector✓ comp disclosed14 days ago
Director-level PM building Claude-powered security products; owns model capability gaps and shipping strategy.
eval designmodel selectionhallucination uxhuman-in-the-loopai safety policyexperimentation
- Claude is now good at security work. It finds and fixes vulnerabilities in real codebases, triages alerts
- Know where the model falls short on security tasks and help build a plan to improve
- You treat what the model can't do yet as a research target with a reliability bar, and you know what to ship in the meantime
- Fluency with AI products: evaluations, model behavior as part of the product surface
- Vulnerability discovery and remediation in Claude Code today and a broader set of capabilities in the future
Applied featureLead14 days ago
PM for age-gated AI safety features (parental controls, crisis support) in ChatGPT for teens.
hallucination uxai safety policyhuman-in-the-loop
- builds features such as age prediction, parental controls, and access to crisis support resources
- create age specific product experiences for <18 users
- balance usefulness, simplicity, trust, safety, and accessibility
- no mention of model selection, evals, fine-tuning, latency budgets, or concrete AI artefacts the PM owns
Internal AI opsSenior IC✓ comp disclosed14 days ago
PM shipping Claude-powered internal tools and automation across Anthropic functions.
prompt engineeringcost modellingexperimentationml metricshuman-in-the-loop
- Own the roadmap for Claude-powered internal products and automation
- Stay current on Claude capabilities and bring new ones to internal teams where they fit
- You have built with LLMs or agent tooling (nice to have)
- Partner with engineering on architecture choices that keep internal products maintainable
Internal AI opsSenior IC✓ comp disclosed14 days ago
PM shipping Claude-powered internal tools and automation across Anthropic functions.
prompt engineeringcost modellingexperimentationml metrics
- Own the roadmap for Claude-powered internal products and automation
- Stay current on Claude capabilities and bring new ones to internal teams where they fit
- You have built with LLMs or agent tooling (nice to have)
- Partner with engineering on architecture choices that keep internal products maintainable
Applied featureDirector✓ comp disclosed15 days ago
Director-level PM shipping Claude-powered security products; owns model capability gaps and reliability.
eval designmodel selectionhallucination uxhuman-in-the-loopai safety policy
- Claude is now good at security work. It finds and fixes vulnerabilities in real codebases, triages alerts
- Know where the model falls short on security tasks and help build a plan to improve
- You design for where the models will be rather than where they are. You treat what the model can't do yet as a research target with a reliability bar
- Fluency with AI products: evaluations, model behavior as part of the product surface
- Ship security products. Vulnerability discovery and remediation in Claude Code
Applied featureLead✓ comp disclosed16 days ago
PM shipping Claude into government workflows; owns model/product decisions for public sector use cases.
eval designmodel selectionhuman-in-the-loopexperimentation
- Fluency with AI products: evaluations, model behavior as part of the product surface
- Own product decisions for our government models and product offerings
- shape the product around what you find [in discovery]
- Spot the government needs worth generalizing and carry them into Anthropic's broader roadmap
Applied featureSenior IC16 days ago
Forward-deployed technical PM selling and implementing Claude solutions to Japanese enterprises
prompt engineeringeval designmodel selectionhallucination uxai safety policy
- Develop customized pilots and prototypes, as well as evaluation suites to make the case for customer adoption
- Leveraging novel prompting techniques
- Configure Claude's capabilities to showcase how it meets the customer's needs
- Recent experience building production systems with large language models
Applied featureSenior IC16 days ago
PM for ChatGPT/Codex app ecosystem and partner integrations; strategy and 1P experiences.
prompt engineeringai safety policy
- mentions 'model capabilities' and 'model-behavior insights' as inputs to prioritization
- references 'model/product constraints' as something to reason about
- no concrete ownership of evals, model selection, latency budgets, data labelling, or failure handling
- focus is on app ecosystem strategy and partner launches, not model quality or AI artefacts
AI-native 0→1Staff✓ comp disclosed17 days ago
Staff PM building Harvey's new AI-powered contract intelligence product from 0-to-1 for legal professionals
model selectionprompt engineeringhuman-in-the-loopexperimentationml metrics
- Define and own the product vision for Harvey's Contracts product
- work at the intersection of cutting-edge AI technology and complex legal workflows
- Translate legal domain expertise into product requirements, working with legal engineers and subject matter experts to build AI-powered features
- Strong technical fluency and ability to collaborate deeply with engineering teams on complex AI/ML-powered products
- Experience with AI/ML products, particularly generative AI or large language models
AI-native 0→1Staff✓ comp disclosed17 days ago
Staff PM building Harvey's contract intelligence product from 0-to-1 using frontier agentic AI
model selectionprompt engineeringhallucination uxhuman-in-the-loopexperimentationml metrics
- Lead 0-to-1 product development of AI-powered contract product
- Work at intersection of cutting-edge AI technology and legal workflows
- Translate legal domain expertise into product requirements working with legal engineers
- Strong technical fluency and ability to collaborate deeply with engineering teams on complex AI/ML-powered products
- Experience with AI/ML products, particularly generative AI or large language models (preferred)
Applied featureLead17 days ago
Design leader for B2B growth experiences on Codex/ChatGPT; no AI technical ownership.
experimentationhallucination ux
- Translate complex and rapidly evolving AI capabilities into clear, approachable experiences
- Help define what AI-native growth experiences should look, feel, and behave like
- Are fluent in using AI as part of how you explore, prototype, communicate, and create
- No mention of model selection, evals, latency budgets, data labelling, or any concrete technical AI artefact the PM owns
Applied featureSenior IC18 days ago
Full-stack engineer building UI/backend for persistent AI agent products at OpenAI
human-in-the-loopagent behaviour
- Role is full-stack product engineer, not PM
- Focuses on UX/frontend/backend for agent workflows, not model or eval ownership
- Mentions 'agent capabilities' and 'agentic workflows' but names no concrete technical AI responsibility (no evals, model selection, fine-tuning, data labelling)
- Owns 'user controls, understandable agent behavior, safe execution' but these are product/UX concerns, not AI technical decisions
Model / platformLead✓ comp disclosed20 days ago
Founding PM for Claude Platform billing, monetization, and vertical expansion; owns commercial machinery and go-to-market.
cost modellingexperimentationml metrics
- owns metering and attribution of usage through Claude API
- responsible for spend observability and controls for enterprise buyers
- must understand how billing, identity, or distribution changes when customer is sometimes an agent
- owns outcome metrics including volume unified across API and subscriptions
- however, no mention of model evaluation, fine-tuning, prompt engineering, or model quality ownership
Applied featureLead✓ comp disclosed21 days ago
Engineering Manager leading teams that build and deploy AI agents for customer support; not a PM role.
- AI agents are the product surface, but no concrete technical PM responsibility is named
- mentions 'outperform human agents' and 'complex customer interactions' but no evaluation methodology, metrics, or model quality ownership
- says 'work with product, design and research' and 'help define technical strategy' but the JD is for an Engineering Manager, not a PM
- no mention of evals, model selection, fine-tuning, data labelling, latency budgets, or failure handling design
Applied featureLead✓ comp disclosed21 days ago
Engineering Manager leading agent deployment teams; not a PM role.
- AI agents are the product surface: 'building and shipping best-in-class AI agents'
- No concrete technical artefact ownership named: no mention of evals, model selection, latency budgets, data labelling, or failure handling
- Role is engineering management, not product management: 'lead a group of engineers', 'mentoring engineers', 'technical guidance'
- Customer-facing but no PM-specific AI responsibility: 'partner directly with enterprise customers' and 'translate them into scalable AI agent solutions' are delivery, not product strategy
Applied featureDirector21 days ago
Senior PMM leading go-to-market strategy and narrative for ChatGPT Work product
- owns ChatGPT Work product marketing strategy and narrative
- translate product capabilities into customer workflows
- shape product through customer and market insight
- no mention of eval design, model selection, fine-tuning, latency budgets, or any concrete technical AI artefact the PM owns
Applied featureStaff21 days ago
Staff PM owning enterprise adoption of Replit Agent coding assistant, governance, and land-expand motion
agent behaviourprompt engineeringhuman-in-the-loop
- Make Replit Agent the best in the world at solving business problems for non-technical users
- Turn our coding agent into the fastest path from idea to solving a real problem through software, with the connectors, skills, and workflows enterprise teams need
- Own the capabilities serious buyers depend on: identity and access (SSO, SCIM), admin and governance, data privacy and compliance, and flexible deployment options
AI-native 0→1Senior IC21 days ago
Senior engineer building 0→1 AI-native products; agent behavior and agentic stack ownership.
agent behaviourprompt engineeringexperimentation
- Understanding of the full agentic software development stack, helping coding agents build, test and review correct code
- Helping Replit Agent build viable alternatives to the world's most popular SaaS vendors
- Replit Agent becoming a full principal within a large number of third party services
Internal AI opsSenior IC21 days ago
Growth marketing IC using AI automation to build self-improving paid acquisition engine; not an AI PM role.
experimentationcost modelling
- says 'AI-native product growth lead' and 'AI-powered workflows' repeatedly
- names 'AI-driven ad generation, testing, and iteration' as a deliverable
- but owns no concrete AI artefact: no eval design, no model selection, no fine-tuning, no data labelling
- the AI is a tool (automation, APIs, workflows) not the product PM owns
- no mention of failure modes, hallucination, latency budgets, or model quality trade-offs
Applied featureSenior IC21 days ago
Data scientist for product analytics at an AI-native platform; not an AI PM role.
experimentationml metricscost modelling
- AI agents mentioned as product surface: 'Own the analytics for core product areas — Growth (activation, engagement, monetization & retention), feature adoption, product quality, AI agent effectiveness'
- No concrete technical ownership of AI artefacts named: no mention of eval design, model selection, fine-tuning, prompt engineering, or failure handling
- Role is product analytics/data science, not AI PM: 'Design and analyze product experiments', 'Develop predictive models', 'Model the aha moment' — all standard DS work
- AI appears as a tool for the analyst's own workflow, not as a product responsibility: 'You leverage AI tools extensively in your own analytical workflow'
Applied featureSenior IC21 days ago
Senior frontend engineer shipping AI-powered visual editing and artifact generation products at Replit
hallucination uxhuman-in-the-loopagent behaviourexperimentation
- Invent and build intuitive interfaces for generating, editing, arranging, and collaborating on AI-created presentations, data applications, and visual artifacts
- Design visual tools that make AI-generated work feel as precise, editable, and controllable as traditionally authored content
- Develop feedback loops and evaluations that improve the visual quality of Agent-generated artifacts
- Understanding of the full agentic software development stack, helping coding agents build, test and review correct code
Applied featureSenior IC21 days ago
Senior PM for AI-powered developer experience at Replit; UX-focused, no model ownership.
hallucination uxprompt engineering
- Design AI-powered features that feel intuitive and magical to users
- Own user journey optimization from onboarding through advanced AI-assisted development
- Deep intuition about AI trends and products, particularly in automation and natural language interfaces
- No mention of evals, model selection, latency budgets, data labelling, or any concrete technical artefact the PM owns
Applied featureSenior IC✓ comp disclosed21 days ago
PM for Duet, an AI agent optimization platform that analyzes production conversations and auto-improves agent workflows.
eval designmodel selectionhuman-in-the-loopagent behaviourml metricsexperimentation
- Develop evaluation criteria for Duet's AI output quality
- engage meaningfully with engineers on model behavior, and form your own opinions about whether an AI output is good or not
- Define the metrics that matter: not just adoption, but whether Duet is actually making agents better over time
- analyzes production conversations, identifies where agents fall short, generates and validates AOP updates
- Experience shipping AI-powered features or working with LLMs in a product context is a strong plus
Applied featureLead✓ comp disclosed21 days ago
PM for enterprise agent platform integrations, deployment, and developer experience at conversational AI company.
eval designexperimentationagent behaviourai safety policy
- owns eval, analytics, testing, experimentation as foundational platform features
- responsible for agent building and versioning infrastructure
- must decide what to build natively vs self-serve vs open for extension
- focus on integrations and deployment models for enterprise AI agents
Applied featureStaff✓ comp disclosed21 days ago
Staff engineer owning AI agent product design, evals, and model integration for enterprise customer experiences
eval designmodel selectionagent behaviourhuman-in-the-loopexperimentation
- Design and build AI agents that outperform human agents in managing complex customer interactions
- Experiment with and run evaluations on the latest text and voice models, then integrate them at scale
- Identify cross-customer trends that guide the evolution of Decagon's agent building platform
- Complete ownership and autonomy in building and shipping best-in-class AI agents, from initial implementation through continuous iteration
Applied featureSenior IC21 days ago
Senior IC PM owning end-to-end AI agent design, deployment, and optimization for Fortune 500 customers.
agent behaviourhuman-in-the-loopprompt engineeringexperimentationml metrics
- Design, build, and optimize enterprise-grade AI agents in production
- going deep on each customer's workflows, pain points, and goals
- own the technical win on Decagon's largest strategic opportunities
- Direct Forward-deployed engineers and partner teams to deliver at scale while you stay accountable for the outcome
- Run tight feedback loops into Product and Engineering, shaping the product based on what you learn on the front lines
Applied featureSenior IC✓ comp disclosed21 days ago
Senior engineer building and evaluating AI agents for enterprise customer support at scale.
eval designmodel selectionagent behaviourhuman-in-the-loop
- Design and build AI agents that outperform human agents
- Experiment with and run evaluations on the latest text and voice models
- integrate them at scale with large enterprise-grade customers
- Identify cross-customer trends that guide the evolution of Decagon's agent building platform
Applied featureSenior IC✓ comp disclosed21 days ago
Senior IC PM building and deploying enterprise AI agents for Fortune 500 customers, owning end-to-end agent design and production optimization.
agent behaviourprompt engineeringhuman-in-the-loopexperimentation
- Design, build, and optimize enterprise-grade AI agents in production
- going deep on each customer's workflows, pain points, and goals
- understand and shape AI agent designs and hold your own in a technical conversation with both engineers and executives
- owns the technical win on Decagon's largest strategic opportunities
Applied featureStaff✓ comp disclosed21 days ago
Staff engineer leading full-stack platform for customers to build and optimize AI agents; light AI depth.
prompt engineeringagent behaviourhuman-in-the-loop
- Create AI-powered tools that enable non-technical teams to create and manage sophisticated workflows
- Build monitoring and analytics that surface actionable insights, helping customers identify where their agents can improve
- Familiarity with LLMs, AI agent systems, prompt engineering, or building products on top of AI capabilities listed as 'even better', not required
- No mention of model selection, evaluation design, fine-tuning, data labelling, or concrete AI artefacts the PM owns
Applied featureSenior IC✓ comp disclosed21 days ago
Senior full-stack engineer building self-serve UI/UX for AI agent configuration and monitoring.
prompt engineering
- says 'AI agents' and 'AI-powered tools' repeatedly but names no concrete technical PM responsibility
- mentions 'prompt engineering' and 'LLM' only in 'even better' section as nice-to-have, not core
- owns 'configuration abstractions' and 'monitoring' but these are generic product work, not AI-specific artefacts
- no mention of evals, model selection, latency budgets, data labelling, or failure handling
Applied featureSenior IC✓ comp disclosed21 days ago
Senior IC PM building and deploying enterprise AI agents for Fortune 500 customers, owning end-to-end agent design and production outcomes.
agent behaviourhuman-in-the-loopprompt engineeringexperimentation
- Design, build, and optimize enterprise-grade AI agents in production
- going deep on each customer's workflows, pain points, and goals
- understand and shape AI agent designs and hold your own in a technical conversation
- owns the technical win on Decagon's largest strategic opportunities
- Direct Forward-deployed engineers and partner teams to deliver at scale while you stay accountable for the outcome
Applied featureSenior IC✓ comp disclosed21 days ago
Senior full-stack engineer building self-serve UI/UX for AI agent configuration, not a PM role.
- says 'AI agents' and 'AI-powered tools' multiple times but names no concrete technical PM responsibility
- mentions 'prompt engineering' and 'LLM' only in 'even better' section as nice-to-have, not core
- describes building 'configuration abstractions' and 'monitoring' but no ownership of model quality, evals, or failure modes
- role is full-stack software engineer, not PM—owns 'features end-to-end' from architecture through deployment, not model outcomes
Applied featureSenior IC✓ comp disclosed21 days ago
Senior IC PM building and deploying enterprise AI agents for Fortune 500 customers at Decagon.
agent behaviourhuman-in-the-loopprompt engineeringexperimentation
- Design, build, and optimize enterprise-grade AI agents in production
- going deep on each customer's workflows, pain points, and goals
- understand and shape AI agent designs and hold your own in a technical conversation with both engineers and executives
- Direct Forward-deployed engineers and partner teams to deliver at scale while you stay accountable for the outcome
Applied featureLead✓ comp disclosed21 days ago
Founding PM for conversational AI agent platform core capabilities (testing, analytics, configuration).
experimentationagent behaviourhuman-in-the-loopml metrics
- owns agent building, testing, experimentation, analytics, QA, versioning
- designing AI-powered features that turn conversation data into actionable insights
- building intelligent systems that give enterprises control over how agents learn, adapt, and deploy
- creating self-serve configuration experiences powered by real-time AI guidance
Applied featureSenior IC✓ comp disclosed21 days ago
Senior engineer building and evaluating AI agents for enterprise customer support at scale.
eval designmodel selectionagent behaviourhuman-in-the-loop
- Design and build AI agents that outperform human agents
- Experiment with and run evaluations on the latest text and voice models
- integrate them at scale with large enterprise-grade customers
- Identify cross-customer trends that guide the evolution of Decagon's agent building platform
Applied featureSenior IC✓ comp disclosed21 days ago
Senior IC PM building and deploying enterprise AI agents for Fortune 500 customers at Decagon.
agent behaviourprompt engineeringhuman-in-the-loopexperimentation
- Design, build, and optimize enterprise-grade AI agents in production
- going deep on each customer's workflows, pain points, and goals
- understand and shape AI agent designs and hold your own in a technical conversation with both engineers and executives
- Turn the frameworks and playbooks you develop into reusable assets
Internal AI opsSenior IC✓ comp disclosed21 days ago
Own evaluation quality and human-in-the-loop data ops for Harvey's legal AI platform at scale.
eval designhuman-in-the-loopdata labellingml metricsexperimentation
- Own the quality bar for Harvey's human evaluations: define what 'good' looks like for eval methodology and data analysis
- Author and maintain the evaluation guidelines, instructions, and databases
- Own contract-attorney quality: onboarding, calibration training, inter-rater reliability, and the feedback loop
- Conducting quantitative and qualitative statistical analyses, diagnosing error states, investigating root causes
- Experience with running quantitative and qualitative data analyses, interpreting and running statistical tests on evaluation data
- Built calibration, inter-rater reliability, or capability-tracking systems for annotation or evaluation pipelines
- Familiarity with LLM-as-judge / automated evaluation used alongside human eval
Applied featureSenior IC21 days ago
Forward-deployed PM embedding AI workflows at law firms; client-facing innovation role, not model ownership.
prompt engineeringhuman-in-the-loopagent behaviour
- describes role as 'blends elements of product management, innovation strategy, solutions design, and forward-deployed partnership'
- mentions 'translate messy, high-value legal work into AI-powered workflows'
- says 'work with innovation leaders' and 'shape how the legal industry adapts to AI'
- no mention of evals, model selection, latency budgets, data labelling, or any concrete technical artefact the PM owns
- focus is on client partnership, workflow design, and go-to-market execution, not model quality or AI system ownership
AI infraStaff✓ comp disclosed21 days ago
Staff PM owning infrastructure and platform reliability for enterprise legal AI system at scale
latency budgetingcost modellingml metricsexperimentation
- owns infrastructure planning and roadmapping for platform processing trillions of tokens and millions of daily requests
- defines metrics for adoption, engagement, quality, and business impact
- projects include architecting multi-region deployment strategies and developing comprehensive observability infrastructure
- background with LLM-powered applications and high-throughput inference systems listed as nice-to-have
Applied featureStaff✓ comp disclosed21 days ago
Staff PM for Vault, Harvey's AI-powered document/knowledge management platform powering legal AI workflows
rag vs finetunemodel selectioncost modellingexperimentationml metrics
- AI-enabled search and Q&A
- document extraction
- Experience with AI/ML-powered products, semantic search, retrieval systems, or document processing pipelines (preferred)
- retrieval architectures (technical acumen requirement)
Applied featureStaff✓ comp disclosed21 days ago
Staff PM building legal operations admin platform on top of Harvey's AI foundation; no model ownership.
experimentationml metrics
- Role owns legal operations platform features, not AI model or evaluation
- Mentions 'AI-native software reshapes a decades-old legal operations workflow' but no concrete AI artefact ownership
- Requires 'strong technical acumen' and 'retrieval architectures' but no mention of evals, model selection, fine-tuning, or failure handling
- Focus is on 'configure and govern relationships with outside firms' and 'performance visibility' — operational tooling, not model quality
Applied featureStaff✓ comp disclosed21 days ago
Staff PM for Vault, Harvey's AI-powered document/knowledge management platform for legal professionals
rag vs finetunemodel selectioncost modellingexperimentationml metrics
- AI-enabled search and Q&A
- document extraction
- Experience with AI/ML-powered products, semantic search, retrieval systems, or document processing pipelines (preferred)
- retrieval architectures (technical acumen requirement)
Applied featureStaff✓ comp disclosed21 days ago
Staff PM for legal operations admin platform built on Harvey's AI backbone; no direct AI model ownership.
- describes building 'legal operations platform' and 'admin experience for in-house legal teams' with no mention of model evaluation, fine-tuning, or concrete AI artefacts
- says 'AI-native software reshapes a decades-old legal operations workflow' but owns workflow/operations, not AI quality
- asks for 'strong technical acumen' and 'retrieval architectures' but these are infrastructure concerns, not PM ownership of model behaviour
- no mention of evals, model selection, hallucination handling, or failure modes
Applied featureStaff✓ comp disclosed21 days ago
Staff PM for legal operations admin platform at AI-native company; enterprise SaaS, not AI-specific PM work
- owns legal operations platform features, not AI model or evaluation
- mentions 'AI-native software reshapes workflow' but no concrete AI artefact ownership named
- no mention of evals, model selection, latency budgets, data labelling, or failure handling
- focus is enterprise SaaS platform PM (config, governance, compliance, visibility) with AI as context, not substance
Applied featureSenior IC✓ comp disclosed21 days ago
Product Data Scientist defining metrics and running experiments for an AI-native legal services platform.
experimentationml metricscost modelling
- Role is explicitly 'Product Data Scientist' not PM; focuses on metrics, experimentation, and analytics
- Mentions 'measure the impact of product, model, and workflow changes' but owns no model decisions
- Lists 'AI/ML products, large language models' as bonus experience, not core responsibility
- No mention of eval design, model selection, fine-tuning, hallucination handling, or agent behaviour
- Core work is measurement and experimentation, not AI product ownership
Internal AI opsSenior IC✓ comp disclosed21 days ago
GTM ops PM building AI agents and automations for post-sales customer lifecycle at Harvey
human-in-the-loopagent behaviourprompt engineeringdata labellingcost modelling
- Leverage applications (Ticketing, CSP, PSA, etc.) and AI agents to automate customer lifecycle
- Work closely with managers and ICs across functions to provide expert level advice about what AI is ready to do, and what should be left to humans
- Ability to demonstrate a portfolio of AI first principles. You will be expected to both explain the technical nature of different models and platforms, and show micro-apps you have built on Claude, Lovable, n8n, Replit, or other agentic applications
- Deep focus on the orchestration and data layers
Internal AI opsSenior IC✓ comp disclosed21 days ago
GTM ops PM building AI agents and automations for post-sales customer lifecycle efficiency
human-in-the-loopagent behaviourprompt engineeringdata labellingcost modelling
- Leverage applications (Ticketing, CSP, PSA, etc.) and AI agents to automate customer lifecycle
- Work closely with managers and ICs across functions to provide expert level advice about what AI is ready to do, and what should be left to humans
- Ability to demonstrate a portfolio of AI first principles. You will be expected to both explain the technical nature of different models and platforms, and show micro-apps you have built on Claude, Lovable, n8n, Replit, or other agentic applications
- Deep focus on the orchestration and data layers
Internal AI opsSenior IC✓ comp disclosed21 days ago
GTM ops PM building internal agentic automation for post-sales lifecycle; owns orchestration and human-AI decisions.
agent behaviourhuman-in-the-loopdata labellingcost modelling
- Leverage applications (Ticketing, CSP, PSA, etc.) and AI agents to automate customer lifecycle
- Work closely with managers and ICs across functions to provide expert level advice about what AI is ready to do, and what should be left to humans
- Ability to demonstrate a portfolio of AI first principles. You will be expected to both explain the technical nature of different models and platforms, and show micro-apps you have built on Claude, Lovable, n8n, Replit, or other agentic applications
- Deep focus on the orchestration and data layers
Internal AI opsSenior IC✓ comp disclosed21 days ago
GTM ops PM building AI agents and automations for post-sales customer lifecycle at Harvey
human-in-the-loopagent behaviourprompt engineeringdata labellingcost modelling
- Leverage applications (Ticketing, CSP, PSA, etc.) and AI agents to automate customer lifecycle
- Work closely with managers and ICs across functions to provide expert level advice about what AI is ready to do, and what should be left to humans
- Ability to demonstrate a portfolio of AI first principles. You will be expected to both explain the technical nature of different models and platforms, and show micro-apps you have built on Claude, Lovable, n8n, Replit, or other agentic applications
- Deep focus on the orchestration and data layers
Internal AI opsSenior IC✓ comp disclosed21 days ago
GTM ops PM building internal agentic automation for post-sales lifecycle; owns orchestration and human-AI decision boundaries.
human-in-the-loopagent behaviourdata labellingcost modelling
- Leverage applications (Ticketing, CSP, PSA, etc.) and AI agents to automate customer lifecycle
- Work closely with managers and ICs across functions to provide expert level advice about what AI is ready to do, and what should be left to humans
- Ability to demonstrate a portfolio of AI first principles. You will be expected to both explain the technical nature of different models and platforms, and show micro-apps you have built on Claude, Lovable, n8n, Replit, or other agentic applications
- Deep focus on the orchestration and data layers
AI infraSenior IC21 days ago
Build evaluation infrastructure and ops systems for AI model quality at Harvey
eval designml metricsexperimentationdata labelling
- Build and scale the systems that power model and product evaluations
- Create the single source of truth for evaluation status, results, history, and launch readiness
- Turn Expert-designed evaluation methodologies into scalable, repeatable operational processes
- Drive evaluation readiness for major product and model launches across geographies and jurisdictions
- ensuring our models behave reliably, accurately, and jurisdictionally correctly is mission-critical
- Experience working with ML/AI evaluations, benchmarking frameworks, or scientific workflows
AI-native 0→1Staff✓ comp disclosed21 days ago
Staff PM building new AI-native product at Harvey; owns 0→1 strategy for legal AI platform integrations and ecosystem.
rag vs finetunecost modellingdata labellingexperimentationml metrics
- owns product strategy at intersection of AI, legal data, and legal work
- work with integrations engineering pod on data pipelines and retrieval architectures
- define success metrics for health and iterate based on data
- strong technical acumen with system design, distributed systems, data pipelines
AI-native 0→1Director✓ comp disclosed21 days ago
Director of Engineering leading Harvey's core product platform; engineering leadership role, not AI PM
- says 'frontier agentic AI' and 'AI-native legal applications' but names no concrete technical PM responsibility
- mentions 'nuances of AI product development' as a requirement but JD itself describes no eval, model selection, latency budget, or data labelling ownership
- role is primarily engineering leadership (hiring, org scaling, platform architecture) not AI product management
Applied featureStaff✓ comp disclosed21 days ago
Staff PM scaling AI agent solutions across enterprise verticals, turning bespoke deployments into repeatable products.
agent behaviourhuman-in-the-loopml metricsexperimentation
- Map workflows to the broader product, identifying where out-of-the-box agent capabilities fit and where custom development is needed
- Define and track success metrics for the verticals you support, ensuring solutions deliver measurable productivity and quality outcomes
- Experience building or deploying AI agents, workflow automation tools, or retrieval-augmented systems (bonus)
- translate those into agent-powered solutions (out-of-the-box or net-new)
Applied featureSenior IC✓ comp disclosed21 days ago
Forward-deployed PM embedding AI workflows into law firm operations; shapes pilots and rollouts with clients.
human-in-the-loopprompt engineeringexperimentation
- translate messy, high-value legal work into AI-powered workflows
- workflow design, testing, onboarding, iteration
- conversations about generative AI, workflow design, privacy, governance, and enterprise rollout
- Help clients evaluate the business impact of AI — from improving associate leverage to transforming workflows
Applied featureStaff✓ comp disclosed21 days ago
Staff Product Designer for AI-powered legal platform; UX for AI outputs, not AI PM
hallucination uxhuman-in-the-loop
- crafting interfaces that help users understand, trust, and effectively guide AI outputs
- Promote and develop best practices in designing with AI
- no mention of eval design, model selection, latency budgets, data labelling, or any concrete AI artefact the designer owns
Applied featureStaff✓ comp disclosed21 days ago
Staff PM owning integrations and data connectors for agentic legal AI platform; light AI depth.
prompt engineering
- says 'frontier agentic AI' and 'applied AI' multiple times but names no concrete technical responsibility
- mentions 'work with engineers' on 'reasoning, summarization, and legal research workflows' but owns integrations/partnerships, not model quality
- no mention of evals, model selection, latency budgets, data labelling, or failure handling
- role is 'equal parts product intuition and technical execution' on integrations, not AI artefacts
Applied featureLead21 days ago
PM shipping enterprise AI agents for customer service; owns agent behavior and customer fit.
agent behaviourhuman-in-the-loopexperimentation
- build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate our customers' needs
- experience with product development for AI agents a plus
- no mention of eval design, model selection, fine-tuning, or failure handling specifics
Applied featuresenior21 days ago
Senior PM owning production AI agent quality, evals, and escalation logic for flagship enterprise customer at scale.
eval designhallucination uxhuman-in-the-loopagent behaviourml metricsexperimentation
- Own the agent, end to end. You own what the agent does (workflows, decisions, escalation logic), how it works (architecture, data flows, APIs, integrations), and how it feels to users (clarity, trust, edge cases, failures at scale)
- You own the customer experience end to end — running evaluations, digging into failures, identifying the highest leverage experience improvements
- You've worked with language models and production AI agents, and understand not just what they do but also how they work under the hood
- You've shipped and owned production-grade systems at scale, with the reliability, observability, and hard tradeoffs that come with complex systems
Applied featureLead21 days ago
PM shipping AI agent features for enterprise customer service platform; customer-facing, not model-focused.
agent behaviourhuman-in-the-loopexperimentation
- build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate our customers' needs
- experience with product development for AI agents a plus
- no mention of eval design, model selection, fine-tuning, or failure handling specifics
AI-native 0→1Lead21 days ago
PM for Ghostwriter, an agent-building agent; owns end-to-end UX, model selection, eval infrastructure, and human-in-the-loop design.
model selectionhuman-in-the-loopagent behavioureval designhallucination uxprompt engineering
- owns end-to-end product experience from prompt → agent → outcome
- define how AI augments agent development — shape workflows for drafting journeys, run simulations, analyze conversations, improve agents using natural language
- balance autonomy and control — design human-in-the-loop patterns (approval flows, change review, workspace isolation)
- partner on model selection, harness engineering, execution architecture, and evaluation/testing infrastructure
- deeply understand frontier model capabilities and limitations, evaluation challenges, and UX of non-deterministic systems
- zero-to-one and one-to-many role building agent-building agent product
AI infraLead21 days ago
PM for Sierra's agent execution infrastructure platform; owns compute, storage, orchestration, latency and reliability.
latency budgetingcost modellingml metrics
- Define the infrastructure platform: Shape the abstractions that power agent execution—including compute, storage, orchestration, and networking layers
- Own performance, reliability, and scale: Define how Sierra systems handle real-time workloads, high concurrency, and enterprise-grade uptime requirements. Set the bar for latency, availability, and fault tolerance.
- Familiarity with modern AI/ML infrastructure or real-time systems
AI-native 0→1Senior IC21 days ago
PM for Agent Studio: zero-to-one platform defining how teams build, test, and improve AI agents end-to-end
eval designagent behaviourhuman-in-the-loopexperimentationml metrics
- Build simulation and testing systems - Define how agents are validated before deployment. Create tools for simulating real-world scenarios, identifying failures, and improving performance
- Make agent quality measurable and actionable - Define evaluation frameworks and feedback systems so users can understand performance and systematically improve their agents
- Own the Agent Development Life Cycle - Design how users analyze conversations, make changes, test improvements, and release updates
- Integrate AI copilots into the workflow - Decide what is automated vs user-driven, and how AI augments each step
- Experience working with AI systems - Familiarity with LLMs and the challenges of building, testing, and iterating on non-deterministic systems
Applied featureSenior IC21 days ago
PM for real-time voice AI agents; owns interaction model, latency budgets, eval design, and model orchestration.
latency budgetingmodel selectioneval designhuman-in-the-loopagent behaviourml metrics
- Define how we evaluate voice agents: latency, interruption handling, resolution rate, and conversation quality
- Work closely with engineering on streaming architectures, latency budgets, and failure handling
- Partner across ASR, TTS, LLMs, and telephony integrations to deliver a cohesive product. Help decide model choices, orchestration strategies
- Design what 'human-quality' voice interaction actually means in practice
- Build feedback loops that improve performance over time
Applied featureSenior IC21 days ago
Senior PM shipping AI agents for healthcare use cases; owns agent behavior, compliance, and customer experience.
agent behaviourhuman-in-the-loophallucination uxai safety policymodel selectionprompt engineering
- Build enterprise-grade AI agents for healthcare: responsible for partnering with engineers and customers to build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate our customers' needs and business processes
- Shape the platform and healthcare roadmap: surface unmet needs, prototype new tools and features, and collaborate with research, product, and platform to shape the future of AI agent development
- Experiment with the latest voice models and integrate them at scale for healthcare enterprises requiring secure, compliant phone interactions
- Build AI agents for major health insurance networks that handle detailed questions like 'what's my co-pay for a primary care visit?'
- Experience with product development for AI agents a plus
AI-native 0→1Senior IC21 days ago
PM for Sierra's Agent SDK platform; build developer tools for enterprise AI agents, no concrete model ownership named.
agent behaviourhuman-in-the-loop
- says 'AI Agents' and 'agentic systems' repeatedly but names no concrete technical artefact the PM owns
- mentions 'discovering new tools and processes for building agents' without specifying eval design, model selection, or failure handling
- lists 'experience building AI-driven products, especially conversational or agentic systems' as a nice-to-have but does not describe what the PM will decide about agent behaviour, evals, or model quality
- focuses on stakeholder needs and end-to-end ownership but does not name model quality outcomes, latency budgets, or data labelling responsibilities
Applied featureLead21 days ago
PM shipping enterprise AI agents; owns customer fit and agent behavior, not model training.
agent behaviourhuman-in-the-loopexperimentation
- Build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate customers' needs
- Experience with product development for AI agents a plus
- Ability to communicate highly technical concepts including recent AI developments
Applied featureLead21 days ago
PM shipping enterprise AI agents for customer service; owns agent behavior and customer fit.
agent behaviourhuman-in-the-loopexperimentation
- Build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate customers' needs
- AI-related experience (experience with product development for AI agents a plus)
- Ability to communicate highly technical concepts to both non-technical and technically proficient audiences, including recent AI developments
Applied featureLead21 days ago
PM shipping AI agent features for enterprise customer service at Sierra, London-based.
agent behaviourhuman-in-the-loopexperimentation
- build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate our customers' needs
- experience with product development for AI agents a plus
- no mention of eval design, model selection, fine-tuning, or failure handling specifics
Applied featureLead21 days ago
PM shipping enterprise AI agents at Sierra; owns agent quality and customer fit.
agent behaviourhuman-in-the-loopexperimentation
- Build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate customers' needs
- Experience with product development for AI agents a plus
- Ability to communicate highly technical concepts including recent AI developments
Applied featureLead21 days ago
PM shipping AI agent features into Sierra's enterprise customer experience platform
agent behaviourhuman-in-the-loopexperimentation
- build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate our customers' needs
- experience with product development for AI agents a plus
- no mention of eval design, model selection, fine-tuning, or failure handling specifics
Applied featureLead21 days ago
PM shipping AI agent features for enterprise customer service platform
agent behaviourhuman-in-the-loopexperimentation
- Build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate customers' needs
- Experience with product development for AI agents a plus
- No mention of eval design, model selection, fine-tuning, or failure handling specifics
Applied featureLead21 days ago
PM shipping enterprise AI agents for customer service; owns agent behavior and customer fit.
agent behaviourhuman-in-the-loopexperimentation
- Build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate customers' needs
- Experience with product development for AI agents a plus
- Address and overcome technical challenges in the business process
Applied featureLead21 days ago
PM shipping AI agent features into Sierra's enterprise customer experience platform.
agent behaviourhuman-in-the-loop
- build and ship AI agents that handle thousands of customer conversations a day
- Develop and improve Sierra virtual agents to fit and anticipate our customers' needs
- AI-related experience (experience with product development for AI agents a plus)
- no mention of evals, model selection, latency budgets, data labelling, or failure handling
Applied featureSenior IC21 days ago
PM for Sierra's AI agent platform, shipping conversational AI features to enterprise customers.
human-in-the-loopagent behaviour
- says 'AI Agents' and 'AI-driven products' repeatedly but names no concrete technical responsibility
- mentions 'designing for AI-driven products or conversational interfaces' as a nice-to-have, not a core requirement
- no mention of evals, model selection, latency budgets, data labelling, or any artefact the PM owns
- focuses on 'building agents' and 'customer experience' without specifying what PM decisions drive agent quality
Applied featureLead21 days ago
Product marketer scaling ElevenAgents into government/public sector with technical depth on conversational AI.
eval designagent behaviourlatency budgeting
- Technical enough to become our internal expert on conversational agents: latency, telephony, tool calls, evals, and guardrails
- Position our solution for accessibility, compliance, language requirements, and the security and data residency mandates
- Names evals and guardrails but in context of marketing positioning, not PM ownership of model quality
Applied featureSenior IC21 days ago
Manage production teams delivering AI-dubbed content in specific languages; quality & creative direction, not AI PM.
- Role is managing human production teams using ElevenLabs AI audio tools as the delivery mechanism
- Responsibilities focus on onboarding linguists, quality control, and creative direction—not on model behavior, evaluation, or technical AI decisions
- Says 'working with the Productions product and engineering team to optimize our tools' but names no concrete AI artefact the PM owns (no evals, no model selection, no latency budgets)
- AI is the surface (dubbing, audiobooks via AI voices) but the JD describes operational/creative management of human teams, not AI product ownership
AI-native 0→1Lead21 days ago
Lead eng team building security primitives for autonomous AI agents in production codebases
agent behaviourhuman-in-the-loopai safety policyprompt engineering
- owns security primitives that enable agent autonomy and determine what product can ship
- designs agent identity, delegated authority, policy enforcement, secure tool execution, and prompt-injection containment
- must understand how agents use tools, where authority should come from, how prompt injection and exfiltration show up in practice
- responsible for failure handling and human oversight that scales with always-on autonomy
- security outcomes have crisp success metrics tied to agent capability
AI-native 0→1Senior IC21 days ago
Senior product engineer shipping AI-powered coding features; owns agent quality and UX for AI-generated code.
agent behaviourexperimentationhallucination uxprompt engineering
- inventing new interfaces and UX for reviewing PRs of AI-generated code
- running experiments and A/B tests on millions of users to push the frontier of agent quality
- mentions sub-agents, memories, and PR retrieval as concrete technical concerns
- blend excellent engineering with a taste for models and design
Applied featureSenior IC21 days ago
Senior engineer shipping AI-powered features (agents, NLI, suggestions) into Linear's core product.
prompt engineeringfine tuningml metricsagent behaviourhuman-in-the-loop
- Optimize prompts, fine-tune model behavior, and evaluate performance
- Design backend services to power natural language interfaces, smart suggestions, agentic workloads
- Work with product and design to prototype and iterate on intelligent workflows
Applied featureSenior IC21 days ago
Senior full-stack engineer building AI features into Linear's product development platform
- Build AI-powered functionality into the core of Linear
- Agentic workflow system utilizing temporal, operating at production scale
- No mention of model selection, evals, fine-tuning, data labelling, or any concrete AI artefact the engineer owns
Applied featureSenior IC21 days ago
Senior full-stack engineer building AI features into Linear's product development platform
- Build AI-powered functionality into the core of Linear
- Agentic workload system utilizing temporal, operating at production scale
- No mention of model selection, evals, fine-tuning, data labelling, or any concrete AI artefact ownership
Applied featureSenior IC21 days ago
Senior full-stack engineer building AI features into Linear's product dev platform
- Build AI-powered functionality into the core of Linear
- Agentic workload system utilizing temporal, operating at production scale
- No mention of model selection, evals, fine-tuning, data labelling, or any concrete AI artefact the PM/engineer owns
AI infraSenior IC✓ comp disclosed21 days ago
Infrastructure engineer building product analytics and experimentation platforms that serve AI agents at Notion.
experimentationml metrics
- build analytics surfaces that interact with AI agents
- build self-serve analytic platforms that builds on top of events, metrics, and experiments to power workflows used by AI agents
- You've worked on developer or agent-facing tooling, building APIs, MCPs, CLIs
- no mention of model selection, evals, fine-tuning, hallucination handling, or concrete AI technical ownership
Model / platformStaff21 days ago
Senior product marketing leader for OpenAI's API and model portfolio; shapes positioning and go-to-market strategy.
- owns marketing for 'flagship frontier models, specialized models, domain models, and multimodal capabilities'
- must 'understand frontier AI systems, model families, multimodal capabilities, evaluations, tradeoffs, APIs'
- will 'develop early understanding of model capabilities, evaluations, tradeoffs, and technical milestones'
- however: zero mention of owning eval design, model selection decisions, data labelling, latency budgets, cost modelling, or any concrete technical artefact the PM controls
- role is translating and positioning existing models to market, not owning model quality outcomes or technical decisions
Internal AI opsSenior IC21 days ago
Senior full-stack engineer building internal AI-powered workflows for Finance/Supply Chain at OpenAI
agent behaviourhuman-in-the-loopexperimentationml metrics
- Apply AI models, tool use, and agents to business workflows where they create real value, with clear validation, evaluation, human review, and operational safeguards.
- Build durable workflows and agentic systems using Temporal or similar workflow technologies, with appropriate handling for retries, idempotency, human approval, audit logs, observability, replay, and backfill.
- Experience with Temporal, queues, event-driven systems, MCP connectors, AI agents, or evaluation systems is a plus
Internal AI opsLead21 days ago
PM for OpenAI internal IT systems and employee tools, with optional AI/agentic automation focus.
- Use OpenAI models and existing internal capabilities to simplify processes, reduce toil, and increase autonomy and efficiency
- excited to build on OpenAI models and internal agentic platforms, and to use our own technology to improve how we operate
- Experience applying AI and agentic technologies to build intelligent systems that automate workflows
- No mention of eval design, model selection, data labelling, latency budgeting, or any concrete AI artefact ownership
Applied featureLead21 days ago
PM leading OpenAI's legal industry product strategy and execution across teams.
model selectionexperimentation
- reason from model and evaluation evidence
- distinguish model, data, workflow, trust, and delivery problems
- no concrete ownership of evals, fine-tuning, latency budgets, or model quality artefacts named
Applied featureSenior IC21 days ago
PM building AI-powered cyber defense products; owns eval design, safety guardrails, and model-to-product translation
eval designhuman-in-the-loopagent behaviourhallucination uxai safety policyml metrics
- Define practical measures of product quality, including accuracy, usefulness, safety, adoption, and improvements in analyst efficiency
- Help establish appropriate safeguards for permissions, approvals, auditability, tenant isolation, and high-impact security actions
- Familiarity with model evaluations, benchmark design, or human-in-the-loop product workflows
- Establish productive feedback loops between practitioners, product development, evaluations, and model research
- Partner with research and engineering teams to connect model improvements with real product use cases
- An example central measure of progress will be: Measured improvement in defensive outcomes per analyst-hour
Model / platformStaff21 days ago
Staff PM owning model quality, evals, and training decisions for OpenAI's frontier models
eval designmodel selectiondata labellinglatency budgetinghuman-in-the-loopfine tuningml metricsexperimentationcost modelling
- owns model requirements and research priorities across training, inference, and evaluation
- builds closed learning loops turning product usage into datasets, evaluations, experiments, training priorities
- defines success across offline evaluations and online metrics, balancing model quality, latency, safety, cost
- partners on post-training research and integration into mainline model stack
- creates reusable platforms for evaluation, experimentation, and signal collection
- uses product failures to identify gaps and shape research investment
Applied featureStaff21 days ago
PM building security controls and partner interfaces for Codex agent execution and tool use.
ai safety policyhuman-in-the-loopagent behavioureval designhallucination ux
- owns evaluation and launch gates: 'Work with security, safety, research, and engineering teams to test whether controls work under realistic and adversarial conditions. Evaluate risks such as permission bypass, prompt injection, malicious tools, secret exposure'
- defines human-in-the-loop and failure handling: 'Human and policy-based approvals', 'Pausing activity, revoking access, or requiring reauthorization'
- agent-specific security artefacts: 'prompt-injection and untrusted-content defenses', 'Experience with AI agents, MCP, sandboxed execution, prompt-injection defenses, or agent-security evaluations'
- owns model/agent behaviour constraints: 'how identity, permissions, tools, MCP servers, repositories, secrets, networks, and high-impact actions are governed across Codex products'
Applied featureLead21 days ago
PM building AI-powered shopping features in ChatGPT; commerce + discovery focus, not model ownership
hallucination uxexperimentationml metrics
- AI-powered shopping experiences, AI-native shopping experiences mentioned repeatedly
- Partner with research teams to explore how AI can help users discover, evaluate, and purchase
- No concrete technical artefacts named: no mention of evals, model selection, latency budgets, data labelling, or failure handling
- Role is about product vision and cross-functional coordination, not model quality ownership
AI infraLead21 days ago
PM owning safety measurement platforms and metrics for frontier models at OpenAI
eval designml metricsai safety policyexperimentation
- owns OpenAI's approach to measuring harm and safeguard efficacy in production
- determine what we measure, where we measure it, and how we measure it
- represent the company's topline safety metric
- establish repeatable processes to integrate cutting-edge AI safety research into safety measurement products
- develop and continuously refine clear, actionable success criteria that effectively capture our ambitions with safety measurement
AI infraLead21 days ago
PM for OpenAI API infrastructure: data governance, access controls, billing, and enterprise safety systems.
cost modellinglatency budgetingai safety policydata labelling
- Develop an enterprise data control plane to securely access and manage enterprise data for LLM use
- Drive product strategy for safety controls, privacy features, access management, and API usage/cost visibility
- Build and improve data governance capabilities: retention, encryption, audit logs, permissions, and lifecycle management
- Oversee usage metering, cost dashboards, alerts, budgeting tools, and predictable API cost experiences
Internal AI opsLead21 days ago
PM owning safety controls and risk measurement for high-stakes AI deployments and agentic workflows
agent behaviourhuman-in-the-loopml metricsai safety policyexperimentation
- Define how we measure residual risk, decision quality, review quality, and mitigation effectiveness
- Build agentic investigation and context-assembly workflows for complex, multi-step misuse patterns
- owns Integrity's product strategy for sensitive deployments...agentic workflows where harm can unfold across many steps
- build the platform layer that makes them safer to deploy [for agentic AI products]
Model / platformSenior IC21 days ago
PM for OpenAI's agentic APIs, defining developer infrastructure and agent-building capabilities
agent behaviourmodel selectionprompt engineeringexperimentation
- Define strategic priorities and roadmap for improving agentic infrastructure
- Partner with research and engineering teams at a technical level to translate priorities into developer products
- Deeply understand problems faced by agent builders and identify opportunities where products and models can make building agents faster, more intuitive, more reliable, and more powerful
- Proven track record of building for developers, with strong intuition for designing clear, flexible APIs and primitives
Model / platformSenior IC21 days ago
Data Scientist defining metrics and running experiments for OpenAI's API platform and B2B products
ml metricsexperimentationlatency budgetingcost modellingeval design
- Define latency/cost guardrails for new features and models
- Design and interpret A/B tests for new model versions
- Translate product learnings into actionable feedback for Research (e.g., failure modes, eval gaps, model response quality)
- Experience connecting offline evals to online product impact
- Experience defining and operationalizing metrics including reliability/latency/cost and safety
Internal AI opsSenior IC21 days ago
Product engineer building LLM-powered internal automation for OpenAI's sales and revenue operations.
prompt engineeringagent behaviourhuman-in-the-loop
- Build high-impact applications and tools that accelerate OpenAI's go-to-market efforts
- Apply OpenAI's models in novel ways to solve real-world customer and internal workflow problems
- Have built or prototyped LLM-powered workflows using the OpenAI API (or similar)
- charter to automate 100% of digital knowledge work in OpenAI's GTM
Applied featureSenior IC21 days ago
Data scientist embedding with product teams to define metrics and run A/B tests on AI-powered products.
experimentationml metrics
- mentions 'model and UX changes' but owns neither model selection nor evaluation
- A/B testing and metrics are standard product analytics, not AI-specific
- nice-to-have: 'experience in NLP, large language models, or generative AI' — optional, not core
- no mention of eval design, fine-tuning, hallucination handling, or model quality ownership
Applied featureSenior IC✓ comp disclosed21 days ago
Fullstack engineer (not PM) building GenAI agent UIs at Databricks; AI is surface-level in JD.
human-in-the-loopagent behaviour
- Create novel, never-seen-before interfaces for GenAI agents that manage complex workflows while keeping the human in the loop
- Proven experience building and shipping end-to-end generative AI products, ideally with a significant UI/UX component
- says AI/GenAI multiple times but names no concrete technical artefact: no evals, model selection, latency budgets, data labelling, or failure handling mechanisms
AI infraStaff21 days ago
Staff PM for Databricks OLTP/governance features; AI is a customer use case, not core PM responsibility
- mentions 'AI agents' and 'agent-based access via MCP' as a use case
- no concrete AI technical responsibility named (no evals, model selection, fine-tuning, or data labelling)
- role is fundamentally about OLTP databases and data governance, not AI systems
- AI appears as a customer use case, not as the PM's domain of ownership
AI infraStaff✓ comp disclosed21 days ago
Staff PM owning AI platform roadmap (training, serving, vector search, LLMs) at Databricks infrastructure layer
model selectionlatency budgetingcost modellingml metricsexperimentation
- owns product roadmap for AI platform areas — defining what we build, why, and in what order
- drive strategy for key AI platform capabilities, shaping how enterprises operationalize AI at scale
- partner closely with engineering teams to make deeply technical decisions about ML infrastructure — from distributed training architectures to real-time serving systems
- experience with ML/AI infrastructure, data platforms... model training, model serving, feature stores, vector search, LLM infrastructure, ML pipelines
- defines pricing, packaging, and commercialization strategy for AI platform features
Internal AI opsStaff✓ comp disclosed21 days ago
Staff PM building agentic AI automation layer on top of SAP for Databricks internal finance operations
agent behaviourrag vs finetuneeval designhuman-in-the-loopcost modellingml metricsai safety policy
- Design, build, and deploy intelligent agentic workflows on top of SAP, partnering with engineering and data teams on the orchestration, retrieval, and guardrail patterns that make agents safe to act on financial and procurement data
- Hands-on experience or a strong technical understanding of agentic automation, LLM orchestration, retrieval-augmented generation, and evaluation
- Define success metrics (process cycle time, touchless transaction rate, automation accuracy, adoption, close efficiency, cost-to-serve)
- clear judgment about where AI belongs in a controlled, auditable ERP process and where it does not
- designing automation that satisfies governance and audit requirements
AI infraStaff✓ comp disclosed21 days ago
Staff PM owning agentic platform strategy: runtime, evaluation, intelligence layer, MCP connectors, and developer experience.
eval designrag vs finetuneagent behaviourhuman-in-the-loopcost modellinglatency budgetingml metricsai safety policy
- Own the AI-judge evaluation pipeline: offline eval with golden datasets, online LLM-as-judge scoring, domain-specific judges
- Define and drive the agent runtime supporting multi step orchestration with durable execution, model gateway abstraction across all providers, governed tool invocation, and configurable per-agent guardrails
- Establish the intelligence layer. Define the three layer data architecture: knowledge graph, context graph, and temporal memory. Ensure unified retrieval across vector, structured, and graph sources
- Ship the evaluation and quality framework. No agent reaches production without passing quality and safety thresholds
- You know what an agent runtime is, what RAG means in practice, and why evaluation is the hardest part
- Experience defining and shipping developer experiences: SDKs, CLIs, templates, documentation, and self service workflows
AI infraStaff✓ comp disclosed21 days ago
Staff PM owning AI platform roadmap (training, serving, vector search, LLMs) at Databricks infrastructure layer
model selectionlatency budgetingcost modellingml metricsexperimentation
- owns product roadmap for AI platform areas — defining what we build, why, and in what order
- drive strategy for key AI platform capabilities, shaping how enterprises operationalize AI at scale
- partner closely with engineering teams to make deeply technical decisions about ML infrastructure — from distributed training architectures to real-time serving systems
- experience with ML/AI infrastructure, data platforms... model training, model serving, feature stores, vector search, LLM infrastructure, ML pipelines
- defines pricing, packaging, and commercialization strategy for AI platform features
AI infraStaff✓ comp disclosed21 days ago
This is a Staff SRE/Infrastructure role, not a PM role. Not classifiable as AI PM.
- Architecting Agentic Reliability: Define and drive the design of future 'self-healing' infrastructure at scale where AI agents proactively detect, diagnose, and remediate production incidents
- Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks is a significant plus
- leveraging AI/ML to revolutionize infrastructure management
- role is SRE/infrastructure engineering, not product management; AI is a tool/context, not the core PM responsibility
AI infraSenior IC✓ comp disclosed21 days ago
Technical PM for Databricks data/AI platform infrastructure, focusing on OLTP or governance for AI agents.
cost modellinglatency budgetingai safety policyexperimentation
- Define and run performance benchmarks (OLTP focus)
- for governance focus: define processes and mechanisms for how AI agents securely and compliantly access the Databricks Data Intelligence Platform
- enabling governed access for AI-driven use cases (e.g. agent-based access via MCP or similar technologies)
- evaluate how these workloads are implemented on the Databricks Data Intelligence Platform
AI infraSenior IC21 days ago
PM for Databricks' data orchestration platform (Jobs), managing workflows and observability for data/AI teams
experimentationml metrics
- Powers ETL, AI/ML, BI, and streaming workloads
- Evolving control-flow capabilities, authoring experiences, and observability features
- No mention of model evaluation, fine-tuning, prompt engineering, or any concrete AI artefact ownership
- Role is about orchestration platform for data/AI teams, not AI model or feature development
Internal AI opsStaff✓ comp disclosed21 days ago
Staff PM building agentic automation workflows for internal HR/people operations at Databricks
agent behaviourhuman-in-the-loopml metricsexperimentationcost modelling
- owns agentic automation workflows and LLM orchestration design
- defines success metrics including 'automation accuracy'
- responsible for designing intelligent automation layer with human-in-the-loop implications (surfacing tasks/actions)
- hands-on experience with agentic automation and modern AI-driven workflows required
- no mention of model training, fine-tuning, or eval set ownership
AI infraSenior IC✓ comp disclosed21 days ago
Sr PM for Unity Catalog metadata & governance platform; data infra, not AI model PM
- Unity Catalog provides 'governance and security capabilities' for 'data and AI assets'
- Role involves 'defining new platform services' for 'data & AI product teams'
- No mention of model evaluation, fine-tuning, prompt engineering, or concrete AI technical artefacts
- Focus is metadata, governance, lineage, discovery, auditing, compliance—not AI model ownership
AI infraSenior IC✓ comp disclosed21 days ago
Sr PM for Unity Catalog data governance platform; infrastructure PM, not AI model PM
- Unity Catalog provides 'governance and security capabilities' for 'data and AI assets'
- Role involves 'defining new platform services' for 'data & AI product teams' to 'add governance capabilities'
- No mention of model evaluation, fine-tuning, prompt engineering, or concrete AI technical artefacts
- Focus is metadata, security, lineage, discovery, auditing, compliance—data governance infrastructure, not AI model ownership
AI infraSenior IC✓ comp disclosed21 days ago
Sr PM for Databricks Repos source control platform with AI-assisted code features for data/ML teams
prompt engineering
- mentions 'AI native workflow' and 'AI assisted code management features such as automated code suggestions, recommended diffs, merge help'
- no concrete ownership of model evaluation, fine-tuning, data labelling, or model quality metrics
- AI is a surface feature (code suggestions, diffs) but the core role is developer experience and source control workflow
AI infraSenior IC✓ comp disclosed21 days ago
Sr PM for Databricks Repos source control platform with AI-assisted code features
prompt engineering
- mentions 'AI native workflow' and 'AI assisted code management features such as automated code suggestions, recommended diffs, merge help'
- no concrete ownership of model quality, evals, fine-tuning, or any ML artefact
- AI is a surface feature (code suggestions) but the core role is developer experience and source control UX
AI infraSenior IC✓ comp disclosed21 days ago
Sr PM for Databricks' AI platform infra (agents, workflows, orchestration); no model ownership.
- AI is the product surface: 'create foundational capabilities that empower customers to develop agents and models, orchestrate complex workflows'
- No concrete technical artefact ownership named: no mention of evals, model selection, latency budgets, data labelling, or failure handling
- Asks for 'strong technical background in computer science, AI/ML' but role is platform/infra PM, not model quality PM
- Says 'partner with world-class engineering and research teams' but names no specific technical decision the PM owns
AI infraSenior IC✓ comp disclosed21 days ago
Sr PM for Databricks' AI platform infra (agents, workflows, models); shapes enterprise AI strategy, no model ownership.
model selectionexperimentation
- AI is the product surface: 'create foundational capabilities that empower customers to develop agents and models, orchestrate complex workflows'
- No concrete technical artefact ownership named: no mention of evals, model selection, latency budgets, data labelling, or failure handling
- Requires 'strong technical background in computer science, AI/ML' but this is a hiring bar, not a job responsibility
- Role is platform/infra PM ('unified platform that includes Genie, Lakebase, Agent Bricks') not model ownership
Applied featureSenior IC✓ comp disclosed21 days ago
Senior Product Designer for Databricks AI/BI natural language analytics—UX for human-AI collaboration, not PM.
hallucination uxhuman-in-the-loop
- AI appears as product surface: 'design how humans and AI think together', 'AI-generated insights transparent and trustworthy'
- No concrete technical artefacts named: no mention of evals, model selection, latency budgets, data labelling, or failure handling
- Role is design-focused, not PM: 'Product Designer', 'visual and interaction work', 'design patterns for AI-native enterprise software'
- No ownership of model quality outcomes or technical AI decisions
AI infraSenior IC✓ comp disclosed21 days ago
Databricks PM intern across data/AI platform teams; generic PM skills, no AI technical depth required.
- JD mentions 'AI Platform' and 'Machine Learning' as team options but names no concrete technical responsibility
- Says 'you will learn how to be a successful PM' and 'work with engineers' but does not specify ownership of evals, model selection, latency budgets, or data labelling
- Requirement 'You've used AI tooling for both personal productivity and development projects' is surface-level exposure, not technical PM depth
- Impact statements are generic PM work: 'understand customer problem space', 'prototype and test', 'ship features' — no artefact ownership named
Internal AI opsLead✓ comp disclosed21 days ago
GTM enablement lead building AI-powered sales tools and field readiness programs for Databricks product areas.
prompt engineeringagent behaviourrag vs finetune
- Build and ship enablement at scale using AI: use vibe coding, and AI content pipelines to generate first-draft technical deep dives
- Build AI-powered tools that make the field smarter: agents for instant answers, AI role-plays for pitch practice
- Experience building AI applications on modern data and AI platforms (RAG patterns, agent architectures, etc.)
- You've already used AI to build at scale - automating content creation, building internal tools, or shipping demos faster
Applied featureStaff21 days ago
Staff TPM owning AI-powered defense products, with hands-on technical depth in model evaluation and architecture.
eval designmodel selectionexperimentationml metrics
- Design and ship AI-powered products and tooling for defense and national security workflows
- Experience training or evaluating models is a plus
- Deep intellectual curiosity about AI systems — you read papers, dig into technical details
- dig into technical architecture with engineers
- work side-by-side with engineering and ML teams
Applied featureStaff✓ comp disclosed21 days ago
Staff PM building agentic AI platforms for US government national security and defense applications.
agent behaviourprompt engineeringai safety policycost modellinghuman-in-the-loop
- owns end-to-end product development including customer pain points, requirements definition, testing, and launches
- names concrete AI artefacts: agentic applications, Text2SQL intelligence, deep research capability over classified documents
- responsible for 'entire lifecycle of the generative AI platform' including capability prioritization
- experience building 'infrastructure and tooling to develop and support agentic applications' listed as nice-to-have
- however, no mention of eval design, model selection, fine-tuning, or model quality ownership—focus is on platform integration and government deployment
Applied featureSenior IC✓ comp disclosed21 days ago
Senior PM owning RL environments and data strategy for healthcare agent training at Scale AI
eval designdata labellingagent behaviourhuman-in-the-loopml metrics
- own the development of RL environments (the realistic, high-fidelity simulations of healthcare software and workflows that labs use to train and evaluate agents)
- decide what healthcare tasks are worth modeling, how to source and structure the underlying data
- Design and scale Healthcare-specific environments: Scope and build simulations of real Healthcare workflows and environments that let labs train and evaluate agents against real world scenarios
- translate the judgement, edge cases, and workflows you know from the industry into high quality training products
- ML Intuition & Technical Fluency: Enough intuition around how model training and evaluation works and what a Reinforcement Learning environment looks like
Model / platformsenior✓ comp disclosed21 days ago
Senior PM owning Scale's evaluation leaderboard platform, benchmark design, and governance for frontier model measurement.
eval designmodel selectionml metricsexperimentationai safety policy
- Partner with ML researchers and domain experts to develop trustworthy evaluation methodologies, benchmark specifications, and leaderboard scoring frameworks
- Establish governance processes for benchmark quality, evaluation integrity, release management, auditability, and update cadence
- Define and manage the end-to-end leaderboard product lifecycle, from ideation and benchmark design to launch, growth, maintenance, and sunset decisions
- Experience with AI model evaluation, benchmarking, or data products strongly preferred
- Passion for advancing trustworthy AI evaluation and helping define industry standards for measuring frontier model capabilities
Applied featureSenior IC✓ comp disclosed21 days ago
Own Finance vertical for Scale's RL training environments and agent evaluation datasets; design simulations for agent training.
eval designagent behaviourdata labellingml metrics
- own the development of RL environments (the realistic, high-fidelity simulations of financial software and workflows that labs use to train and evaluate agents)
- Design and scale Finance-specific environments: Scope and build simulations of real Finance workflows and environments that let labs train and evaluate agents against real world scenarios
- understand where their Finance agentic capabilities fall short and shape new product lines
- ML Intuition & Technical Fluency: Enough intuition around how model training and evaluation works and what a Reinforcement Learning environment looks like
- Partner with ML and Operations to translate the judgement, edge cases, and workflows you know from the industry into high quality training products
Model / platformsenior✓ comp disclosed21 days ago
Build Scale's cybersecurity eval platform: own task taxonomy, grader design, RL environments, and execution-grounded verification for agentic security models.
eval designmodel selectiondata labellinghuman-in-the-loopagent behaviourfine tuningml metricsai safety policy
- Define the capability map we train and measure against: vulnerability discovery, proof-of-concept reproduction, patch generation and regression safety
- Partner with ML researchers and security practitioners on task specifications, grader design, and verifiable rewards, holding to execution-grounded verification
- a task counts as solved only when the reproducer fires or the patch holds without breaking functionality
- Own the roadmap and strategy for Scale's Cybersecurity portfolio across training data, RL environments, agentic task suites, and evaluation products
- Drive the infrastructure roadmap — reproducible vulnerability images, fuzzing and build toolchains, sandboxed execution, network-segmented ranges, automated verification
- Establish governance for data quality, contamination prevention, license and IP hygiene, reproducibility, and release management
- Familiarity with how models are post-trained and evaluated, including agentic scaffolds and container-based rollout infrastructure
Model / platformsenior✓ comp disclosed21 days ago
Own Scale's coding benchmark and RL environment product line; drive eval design and model training data strategy.
eval designfine tuningagent behaviourdata labellinghuman-in-the-loopml metricsexperimentation
- owns SWE-Bench Pro and SWE Atlas evaluation suites; decides what comes next as agents saturate current tasks
- Partner with ML researchers to develop trustworthy task specifications, rubric and grader design, verifiable reward signals
- Define priorities across SFT and preference data, reinforcement learning environments, agentic task suites
- Establish governance processes for data quality, contamination and leakage prevention
- Track adoption, usage, model-impact signals, and business outcomes
- Familiarity with how coding models are trained and evaluated, including post-training methods, agentic scaffolds and harnesses
Applied featureLead21 days ago
Lead PM building custom LLM applications and AI solutions for government clients, owning eval design and model performance.
eval designmodel selectionfine tuningml metricshuman-in-the-loopexperimentation
- Scope out model evaluation sets and performance requirements, consistently review results, and iterate on the solution
- Lead cross-functional development of AI applications and custom LLMs
- Building custom LLMs
- Stay up to date with latest research in applied AI and training custom LLMs
- owns large AI projects for one or many customers
Applied featureLead21 days ago
Lead PM building bespoke AI applications and custom LLMs for government customers using Scale's platform.
eval designmodel selectionfine tuningml metricshuman-in-the-loopcost modelling
- Scope out model evaluation sets and performance requirements, consistently review results, and iterate on the solution
- Lead cross-functional development of AI applications and custom LLMs
- Building custom LLMs
- Stay up to date with latest research in applied AI and training custom LLMs
- owns large AI projects for one or many customers
AI infraLead✓ comp disclosed21 days ago
Platform PM owning core AI infrastructure (evals, agents, observability) for enterprise AI deployment
eval designfine tuningagent behaviourhuman-in-the-loopml metricslatency budgetingcost modelling
- owns 'agentic primitives, eval infrastructure, expert judgment capture, and data reasoning capabilities that every agent workload needs'
- responsible for defining platform capabilities including 'agent runtime' and 'eval pipelines'
- must understand 'fine-tuning workflows, or observability for production AI systems'
- holds quality bar on capabilities customers 'trust without thinking about it' — implies ownership of failure modes and reliability
- sequences work that unblocks 'hardest customer problems' in agent deployment
Applied featureStaff✓ comp disclosed21 days ago
Forward-deployed PM shipping AI/data products into DoD; owns military planning and real-time alerting outcomes.
latency budgetingagent behaviourhuman-in-the-loopexperimentation
- owns product outcomes for 'cross-cutting portfolio of AI and data capabilities'
- must translate 'real-time, event-driven, high-throughput data architectures and their implications for AI agents'
- nice-to-have includes 'agentic capabilities, human-agent interaction, data labeling, RLHF, fine-tuning workflows, and model evaluation pipelines'
- however, no concrete ownership of evals, model selection, or fine-tuning decisions named as core responsibility
Applied featureLead21 days ago
Forward-deployed PM building bespoke GenAI solutions with enterprise customers using Scale's platform.
prompt engineeringrag vs finetunehallucination uxhuman-in-the-loopexperimentation
- Own end-to-end product development by understanding customer pain points, defining product requirements, managing development, testing, and launches
- Develop enterprise grade solutions that leverage cutting edge AI to drive business value
- 4+ years of experience in building ML-powered products
- Strong understanding of generative AI technologies and their applications in enterprise settings
- Develop a point of view and execute on turning the solutions we build into repeatable software
Applied featureDirector✓ comp disclosed21 days ago
Director leading product strategy and GTM for Scale's data/AI products; no technical AI PM ownership.
- owns 'frontier data products' but no concrete technical artefact named
- partners with 'Researchers and ML Engineers' but PM owns no model quality, eval, or training decision
- domain knowledge in 'Robotics, Autonomous Vehicles, Computer Vision, and/or Machine Learning strongly preferred' but these are hiring preferences, not PM responsibilities
- no mention of evals, model selection, latency budgets, data labelling, or failure handling
- role is portfolio strategy and GTM leadership, not AI technical ownership
Applied featureLead✓ comp disclosed21 days ago
Forward-deployed PM embedding with enterprise customers to drive Scale AI platform deployments to production.
- Preferred qualifications mention 'Direct experience with AI/ML platform products — data labeling, RLHF, fine-tuning workflows, or model evaluation pipelines' but these are optional, not core to the role
- JD focuses on enterprise deployment, customer relationships, and translating requirements—not on owning AI/ML artefacts or outcomes
- No mention of eval design, model selection, latency budgeting, cost modelling, or any concrete technical AI responsibility the PM owns
- Role is fundamentally about field deployment and customer success, not AI product depth
AI infraDirector21 days ago
Director leading forward-deployed PM team converting custom AI data/infra wins into scalable products
data labellingml metrics
- understands how data quality affects model performance
- comfort with APIs, data pipelines, SQL
- high-quality data and full-stack technologies that power the world's leading models
- no mention of eval design, model selection, fine-tuning, latency budgeting, or concrete model quality ownership
AI infraDirector21 days ago
Director leading forward-deployed PM team converting custom AI data/infra wins into scalable products
data labellingml metrics
- understands how data quality affects model performance
- comfort with APIs, data pipelines, SQL
- high-quality data and full-stack technologies that power the world's leading models
- no mention of eval design, model selection, latency budgets, fine-tuning, or concrete technical ownership of model outcomes
Model / platformSenior IC✓ comp disclosed21 days ago
Entry-level AI PM supporting multimodal/coding data products at Scale's data foundry platform.
data labellingeval designmodel selectionml metrics
- assist in defining data specifications, reviewing data quality, and identifying opportunities to improve product performance
- contribute to the development of new AI data products, tooling, and evaluation workflows
- develop a strong understanding of frontier AI models, agent systems, and data pipelines
- support customer engagements by documenting requirements
Internal AI opsStaff✓ comp disclosed21 days ago
Staff engineer building AI-native internal people/HR products at Anthropic using Claude.
eval designprompt engineeringexperimentation
- Design and implement AI-native workflows: build tools, evals, prompts, and products
- Have shipped LLM-native features or applications
- Familiarity with MCP (Model Context Protocol) or prior experience building Claude or LLM integrations in production
Applied featureSenior IC✓ comp disclosed21 days ago
Software engineer shipping Claude Security features; not a PM role.
agent behaviourhuman-in-the-loop
- says 'take new security capabilities in frontier models and turn them into products'
- mentions 'work with researchers on the team to understand what the models can do and where they fall short'
- names 'Build the agent loops, tool integrations' but no concrete PM ownership of model quality, evals, or failure modes
- no mention of eval design, model selection, latency budgets, or data labelling decisions
Applied featureStaff✓ comp disclosed21 days ago
Staff researcher building Claude Security: design evals, measure model capabilities, operationalize for customers
eval designmodel selectiondata labellinghuman-in-the-loopml metrics
- Design evaluations that measure model performance on the work security teams actually do
- Build the datasets, harnesses, and scoring those evaluations depend on
- Track how model capabilities for security are changing, and what that means for what we build next
- identify which security capabilities in frontier models are ready to build on, measure how well they perform
- responsible for answering those questions through fast prototyping and rigorous evaluation
Model / platformLead✓ comp disclosed21 days ago
PM owning Claude's behavioral alignment, evals, and reinforcement signals across model capabilities.
eval designhuman-in-the-loopagent behaviourfine tuningai safety policyml metrics
- Define behavioral defaults and steerability constraints
- Develop and maintain taxonomies of model behaviors across capabilities
- Contribute to evals that measure alignment progress
- Partner with the Alignment Finetuning team to define and shape Claude's character, behaviors, and reinforcement signals
- Amplify alignment research breakthroughs, translating them into product, process, and model improvements
- work that directly influences how millions of people experience AI
AI-native 0→1Lead✓ comp disclosed21 days ago
Lead 0-to-1 product development at Anthropic Labs, turning frontier AI research into new product categories.
prompt engineeringexperimentationmodel selectionhallucination uxagent behaviour
- owns ideation and development of new moonshot products transforming research into applications
- work with researchers to understand emerging capabilities and what it means for users
- identify nascent research capabilities that could become transformative products
- build prototypes yourself to validate ideas
- define product strategy for experimental initiatives leveraging latest AI capabilities
- creatively build MVPs and prototypes to validate product-market fit
- prototype with AI tools like Claude Code
- stay up-to-date and hands-on with emerging research and industry trends
Model / platformSenior IC✓ comp disclosed21 days ago
Research Engineer owning post-training, fine-tuning, and evaluation of production Claude models at scale.
fine tuningeval designml metricshuman-in-the-loopai safety policycost modellinglatency budgeting
- train our base models through the complete post-training stack to deliver the production Claude models
- Implement and optimize post-training techniques at scale on frontier models
- Design, build, and run robust, efficient pipelines for model fine-tuning and evaluation
- Develop tools to measure and improve model performance across various dimensions
- Constitutional AI, RLHF, and other alignment methodologies
- Debug complex issues in training pipelines and model behavior
- experience with training, fine-tuning, or evaluating large language models
Internal AI opsSenior IC✓ comp disclosed21 days ago
Embedded ops PM running model launches, evals, and feedback loops for Anthropic product teams
eval designprompt engineeringmodel selectionexperimentationhuman-in-the-loopai safety policy
- Program manage new model launches on your surface: readiness goals set in advance, testing and eval coverage tracked, issues triaged, prompting changes landed
- Have direct experience managing evals, refining system prompts, and adapting harnesses to new models
- Build with Claude yourself. You have shipped Claude or other LLM-powered workflows, written the prompts, iterated on the outputs, and can describe model behavior with specifics
ai safety policyLead✓ comp disclosed21 days ago
Lead PM for AI safety systems, evals, and risk mitigation across Anthropic's frontier models and products.
eval designhallucination uxhuman-in-the-loopai safety policyml metricsexperimentation
- own the ideation, design, development and deployment of Safeguards systems
- develop detections, evals, interventions, and tools to measure and mitigate deployment and user risks
- Ability to write safety evals and communicate externally about safety
- Lead the development of metrics to understand the area, performance, blindspots
- experience working across policy experts, AI/ML research engineers and software engineering teams to design and build state of the art safety systems
- Demonstrated experience in designing and building metrics to evaluate risks, system performance, user impact
ai safety policyLead✓ comp disclosed21 days ago
Lead PM for AI safety systems, evals, and risk mitigation across Anthropic's frontier models and products.
eval designhallucination uxhuman-in-the-loopai safety policyml metricsexperimentation
- own the ideation, design, development and deployment of Safeguards systems
- develop detections, evals, interventions, and tools to measure and mitigate deployment and user risks
- Ability to write safety evals and communicate externally about safety
- Lead the development of metrics to understand the area, performance, blindspots
- experience working across policy experts, AI/ML research engineers and software engineering teams to design and build state of the art safety systems
- Demonstrated experience in designing and building metrics to evaluate risks, system performance, user impact
Applied featureLead✓ comp disclosed21 days ago
PM owning safety evals, detections, and risk mitigation systems for Claude across deployment surfaces.
eval designhallucination uxai safety policyml metricshuman-in-the-loop
- own the ideation, design, development and deployment of Safeguards systems
- develop detections, evals, interventions, and tools to measure and mitigate deployment and user risks
- Ability to write safety evals and communicate externally about safety
- Lead the development of metrics to understand the area, performance, blindspots
- experience working across policy experts, AI/ML research engineers and software engineering teams to design and build state of the art safety systems
- Demonstrated experience in designing and building metrics to evaluate risks, system performance, user impact
Model / platformLead✓ comp disclosed21 days ago
PM owning frontier model deployment and productization at Anthropic research team
model selectioneval designcost modellingexperimentationml metrics
- own the ideation and deployment of new models and products
- Partner with research to define, improve, and ship model capabilities
- Synthesize user insights into actionable requirements and evaluations
- deep technical background with a strong grasp of AI/ML concepts
- working proficiency in Python and SQL
Model / platformLead✓ comp disclosed21 days ago
Founding PM for Claude Platform billing, monetization, and vertical expansion; owns commercial machinery and go-to-market.
cost modellingexperimentationml metrics
- owns metering and attribution of usage through partners and resellers
- designs spend observability and controls for enterprise buyers
- must have 'built with or shipped on top of large language models, and have a view on what changes about billing, identity, or distribution when the customer is sometimes an agent'
- responsible for platform strategy across verticals and new industries
- no mention of eval design, model selection, fine-tuning, or model quality ownership
Applied featureSenior IC✓ comp disclosed21 days ago
Enterprise PM for Claude.ai security, compliance, and admin features—not an AI-technical role.
- Role is explicitly for 'enterprise platform capabilities' and 'security & compliance' features for Claude.ai
- No mention of model evaluation, fine-tuning, prompt engineering, or any concrete AI technical artefact ownership
- Focus is entirely on enterprise infrastructure, access controls, audit logging, compliance certifications—standard enterprise software PM work
- Mentions 'coordinate with research teams on enterprise-specific model behavior' but PM owns no model quality outcomes
- AI is the product surface (Claude), but the JD names zero technical AI responsibilities
Applied featureSenior IC✓ comp disclosed21 days ago
PM shipping Claude Tag (multiplayer AI agent) to new collaboration surfaces; owns partnerships, adoption, and model behavior tuning.
model selectionagent behaviourhuman-in-the-loopprompt engineering
- debugging a permissions bug or model behavior with engineers
- strong grasp of model capabilities and work fluently with engineering teams on technical products
- Claude Tag writes 65% of merged PRs for our products
- customizing the core interaction paradigm for the surface
- owns adoption from design partners to scaled rollout
Model / platformSenior IC✓ comp disclosed21 days ago
PM owning Claude Code model performance, evals, and launches—direct influence on model behavior and research.
eval designmodel selectionprompt engineeringagent behaviourml metricsexperimentation
- Own model launch planning and execution for Claude Code: define readiness criteria
- Design and implement agentic evals that measure real-world coding performance
- Drive the engineering team's eval roadmap
- Partner with researchers working on coding capabilities to define target behaviors and influence model development with evidence from real usage
- Have personally built agentic evals (e.g. SWE-bench-style task suites)
- comfortable going deep on model behavior, prompt engineering, and evaluation methodology
- you build the infrastructure that prevents its whole class
Model / platformLead✓ comp disclosed21 days ago
PM for frontier AI models at Anthropic, bridging research and product to ship new capabilities.
model selectioneval designcost modellingexperimentationml metrics
- own the ideation and deployment of new models and products
- Partner with research to define, improve, and ship model capabilities
- Synthesize user insights into actionable requirements and evaluations
- deep technical background with a strong grasp of AI/ML concepts
- working proficiency in Python and SQL
AI infraLead✓ comp disclosed21 days ago
PM for Anthropic's data collection and labeling infrastructure platform serving model training.
data labellinghuman-in-the-loopml metricseval design
- Own the product direction for our human data tooling, with clear prioritization across labeling interfaces, infrastructure investments, data quality, and operational visibility
- Define and track outcome-based KPIs: time-to-launch for new data collection projects, end-to-end data quality scores, and measurable impact on model evaluation scores
- Develop a deep understanding of research and training approaches to identify where tooling investments will have the highest leverage
- Experience building data collection tools, annotation platforms, or human-in-the-loop pipelines
- Sit in on crowd worker and vendor sessions to systematically understand pain points
Internal AI opsSenior IC✓ comp disclosed21 days ago
Design AI-native internal HR/people workflows at Anthropic using Claude; no model ownership.
prompt engineeringhuman-in-the-loop
- Design AI-native workflows across the People Products portfolio defining what is possible in applied AI for people processes
- You're AI-native in how you work. You're already using Claude Code, Claude Design, or similar tools to extend what you can build
- You stay close to the models. You pay attention to where capabilities are heading
- No mention of eval design, model selection, data labelling, latency budgeting, or any concrete technical artefact the designer owns
Internal AI opsLead✓ comp disclosed21 days ago
Lead Data Scientist for Anthropic's Developer Platform; analytics and metrics, not product management.
experimentationml metrics
- Role is explicitly a Data Scientist/Analytics position, not a Product Manager role
- Responsibilities focus on metrics, experimentation, and data analysis infrastructure
- No ownership of model quality, eval design, fine-tuning, or concrete AI artefacts
- Mentions 'AI agents being built and deployed' but only as context for analytics, not PM ownership
- Preferred: 'Experience with AI/ML products, large language models' — domain knowledge, not technical PM responsibility
AI infraLead✓ comp disclosed21 days ago
Engineering Manager leading research infrastructure and finetuning pipeline tooling at Anthropic
fine tuningml metricsexperimentation
- builds systems that support large-scale, distributed finetuning runs
- work with research teams to incorporate their innovations into our production finetuning pipeline
- support fast iteration on model development and research
- help us iterate quickly on customer-oriented model improvements
- training runs and data pipelines
Applied featureLead✓ comp disclosed21 days ago
Engineering Manager leading cybersecurity product team shipping LLM-powered defense tools at Anthropic
agent behavioureval designprompt engineeringmodel selectionexperimentation
- Building products that use frontier models to defend code and infrastructure
- Partner with research to turn new model capabilities into products
- Experience with agentic systems, evals, and prompt and model iteration loops
- Own architectural decisions across agentic systems and model orchestration
- No explicit ownership of eval design, fine-tuning, or model quality metrics
Internal AI opsStaff✓ comp disclosed21 days ago
Data scientist measuring developer productivity gains from AI tools; not an AI PM role.
experimentationml metrics
- Role is about measuring developer productivity in an AI-first org, not building AI products
- Mentions 'AI-assisted development' and 'Claude making engineers faster' as measurement targets, not as technical artifacts the PM owns
- No mention of model selection, evals, fine-tuning, latency budgeting, or any concrete AI technical responsibility
- Core work is data science, metrics definition, and experimentation—not AI product management