Available · Seeking North America remote roles

Ali A. Malik

Senior Product Manager, AI

6 years taking AI products from 0 to 1. Now building Proof of Quality at Sapien, the evaluation layer where LLM judges and human experts score AI training data and agent output.

Summary

I'm a senior product manager for AI and AI-assisted platforms, with 6 years taking products from 0 to 1: NLP analytics used by 700,000 people, LLM tooling that cut delivery 40%, and now the evaluation layer frontier-AI teams use to trust their training data.

At Sapien, I own strategy, roadmap, and success metrics for Proof of Quality, a rubric-graded layer where LLM judges and human experts score model and agent output. I have carried it from first spec to live enterprise design partners.

I build the first version in code, run discovery and pilots directly with customers, and keep engineering, design, and go-to-market on one plan.

Experience

Senior Product Manager

Sapien · Full-time

Oct 2025 – Present
Toronto, CA

Human-data and evaluation infrastructure for frontier AI: expert annotation, RLHF, and rubric-based quality attestation.

  • Own Sapien's core product, Proof of Quality: roadmap, evaluation framework, and success metrics for a rubric-graded layer where LLM judges and human experts score training data and agent outputs. Shaped V1 scope and wrote the launch specs.
  • Ran 120+ customer discovery and go-to-market conversations as a forward-deployed PM. Those conversations produced the V1 requirements, 3 pilot proposals for security firms including CertiK and Sherlock, and a converted design partner whose CEO chose Sapien over extending his existing tooling after a live demo.
  • Proved the product with a live enterprise partner by re-evaluating its AI security auditor against 5 senior human auditors: 21 of 40 findings flagged as false positives in 77 minutes, an auditable verdict on 98%, and an approved Phase 2 partnership.
  • Designed and pressure-tested the pricing model in customer conversations, landing on pay-for-the-proof pricing and unit economics ($1,000 per project, per-event fees, $500 per report) so customers pay for verified quality instead of review hours.
  • Resolved 5 launch-blocking defects in the V1 spec before engineering kickoff, including a quality metric defined two ways and AI review panels copying each other's votes. Wrote the rubric guardrails engineering builds against.
  • Built Sapien's internal knowledge brain: a RAG index over all company documents, a Graphify knowledge graph, and a custom agent that answers from it, so planning, strategy, and team updates draw on one accurate source of truth.
  • Ships code alongside engineers: 6 pull requests in the production monorepo, including a 7-layer release gate (UI regression, accessibility, LLM-persona evaluation) that guards every PR. Prototyped the rater marketplace MVP solo.

AI Product Manager

Tempo (Y Combinator 2023) · Full-time

Dec 2024 – Oct 2025
Toronto, CA

Design-to-code AI tooling that turns Figma files into production applications.

  • Shipped 6 production web applications on Tempo's in-house LLM code-generation and agent platform, cutting delivery 40%.
  • Ran 20 concurrent client builds, shipping 1 to 3 features weekly through agile sprints and roadmap ownership.
  • Led 8 developers and 6 designers across fast product cycles, with AI copilots embedded in the design-to-code handoff.
  • Integrated Stripe, BigCommerce, Odoo, and Sanity into modular build pipelines.
  • Built an AI-powered feature tracker that centralized client feedback and gave every build a visible roadmap and delivery status.
  • Led the consumer Kanban initiative that brought AI-assisted project management into daily client work.

Senior Product Manager

Surf · Acquired by Datacy · Full-time

Jan 2023 – Dec 2024
Toronto, CA

NLP and behavioral analytics SaaS; 700,000 users across consumer, business, and marketplace products.

  • Led 2 agile teams across 3 products (consumer, business, and marketplace), supporting 700,000 users with NLP analytics and data intelligence features.
  • Owned the roadmap and KPIs for 34 sprints, shipping 70 features.
  • Grew active users 35% through campaign experiments driven by behavioral analytics and predictive models.
  • Managed enterprise accounts for Amazon, Electronic Arts, Logitech, and Riot Games at $1.5 million average deal size.
  • Built end-to-end data flows for campaign operations, integrating inventory management, OCR receipt validation, purchase tracking, and fulfillment accuracy.
  • Led the stack migration from Elixir and Rails to Node.js: scoped the feature set, wrote the development roadmap, and launched the new release cycles.

Product Manager

Trufan · Acquired by Surf · Full-time

Sep 2020 – Dec 2022
Toronto, CA

Social data and audience intelligence startup.

  • Built and launched 2 consumer platforms serving 500,000 users with a 6-person team: an NLP analytics product over Instagram, Twitter, and web data, and a giveaways platform.
  • Launched a live dashboard aggregating data from 120,000 browser extension users, with SQL-based insights and audience intelligence views.
  • Drove 20,000+ users into the giveaways platform by integrating it with the browser extension's entry flow.
  • Directed 2 major SaaS launches end to end, owning OKRs and cross-functional dependencies.
  • Built data marketplace integrations with Snowflake, LiveRamp, and Eagle Alpha for downstream analytics, segmentation, and activation.

Selected work

Sapien · Proof of Quality

AI / Evaluation

Proof of Quality is Sapien's evaluation layer for AI work: a customer defines quality as a…

Read the case study

artc · Release QA Harness

Eng / QA

artc is Sapien's real-time collaborative document workspace (Rust backend, React frontend, live…

Read the case study

Sapien · Knowledge Brain

AI / RAG

Sapien's institutional memory was scattered across call transcripts, field notes, specs, and…

Read the case study

Sapien · Product Films

Motion / Brand

2 brand films for Proof of Quality, both built entirely in code with Remotion, so every cut is…

Read the case study

mogkit

Open Source

mogkit installs PM craft into the terminal: 13 methodology-backed skills that run in Claude…

Read the case study

Groom Club

eCommerce

Groom Club is a mobile app where dog owners book, manage, and track at-home grooming visits.

Read the case study

Hisense Canada x NBA2K

AI / CV

With Hisense Canada and Surf Giveaways, I led the technical side of a national promotion for the…

Read the case study

App Orchid Insights

NLP / AI

At App Orchid, I coordinated delivery of an NLP dashboard that turned zero-party web browsing…

Read the case study

Sequoia Capital

Data

On a contract with Sequoia's early-stage scouting division, I led a project to surface promising…

Read the case study

Prime Gaming x League of Legends Worlds '22

Scale

For the 2022 League of Legends World Championship, Prime Gaming ran a giveaway across its Twitch…

Read the case study

Novos Fiber

Web

Novos Fiber is a regional high-speed internet provider. The old site was a brochure; the…

Read the case study

Skills

AI evaluation

Evaluation frameworks, rubric and golden-set design, LLM-as-judge with human escalation, inter-rater reliability, RLHF, regression tracking, rollout gates

AI systems

Model selection, LLM and agent workflows, prompt engineering, RAG, MCP, knowledge graphs, guardrails, human oversight, Claude Code, OpenAI

Product

Use-case selection, 0-to-1, roadmaps and OKRs, specifications, success metrics, pricing and unit economics, build-vs-buy, experimentation

Go-to-market

Customer discovery, design-partner pilots, forward-deployed demos, pilot proposals, positioning and messaging, field events, enterprise account management

Building and prototyping

TypeScript, React, Next.js, Node.js, SQL, Snowflake, Playwright, Storybook, GitHub Actions, Supabase, Netlify, Vercel, Remotion

Team and process

Agile and sprint operations, Jira, backlog and release management, QA gates and release readiness, spec reviews, internal knowledge systems, onboarding docs

Leadership

Executive and cross-functional alignment, stakeholder management, client communication, leading engineering and design teams

Domains

AI infrastructure and evaluation for frontier labs, agentic AI, data annotation and computer vision operations, Web3 and blockchain security (Solidity audits, staking), enterprise SaaS, gaming and esports (Riot Games, EA, Prime Gaming), consumer social and audience analytics, eCommerce (Stripe, Odoo), telecom

Speaking

Keynote: training and evaluating AI models

AI Futurist Conference (the AI stage of Blockchain Futurist Conference), Toronto, July 2026.

Independent

Founder of Kleta, a digital agency.

Languages

English (Native)French (Intermediate)Urdu (Conversational)

Education

London School of Economics and Political Science

Bachelor of Science, Business

Interests

VolunteeringYouth MentoringAITravelDigital Nomad LifestyleCasual GamingPower LiftingBook PublishingTypographyMotion DesignEvent OrganizingOpen Source