Skip to proof
Paul Carpenter

AI Product Leadership / PM + Builder

Paul Carpenter

I turn messy, high-stakes workflows into trustworthy systems people can actually ship.

I build AI and product surfaces with review, controls, evidence, rollback paths, and handoff-ready workflows across regulated fintech, eval infrastructure, and internal tools.

failure to reusable eval coverage
EvalGate
regulated scale under release pressure
Chase
lower enrollment time
79%
reduction in production errors
95%
Flagship proof

EvalGate proves the AI thesis. Chase proves the trust under pressure.

The page leads with the work that should carry the hiring signal: AI failures reviewed into reusable coverage, and shipped fintech product work under regulatory, partner, mobile, and customer constraints.

Flagship AI infrastructure productAI PM, developer platform, infrastructure, reliability, and AI-native product roles

EvalGate

AI evaluation control plane for release confidence

  1. 01Failure captured
  2. 02Reviewed case
  3. 03Reusable eval coverage
  4. 04Regression gate

User problemAI teams need a way to capture production failures, turn reviewed failures into reusable eval coverage, and block behavior regressions before prompt, model, retrieval, or provider changes ship.

My roleBuilt and shaped the product surface, SDK workflow, demo data, and infrastructure story around traces, reviewed cases, reusable coverage, judge evidence, CI gates, and release handoffs.

What changedCreated a trust loop for risky AI surfaces: capture the failure, review the case, promote it into eval coverage, gate regressions, and ship only when the evidence clears.

  • AI reliability infrastructure
  • reviewed eval coverage
  • regression gates
  • release confidence
  • auditability
  • handoff-ready evidence
EvalGate dashboard showing evaluations, recent runs, active traces, and average quality score
Control-plane dashboard
EvalGate evaluations list with unit tests, model evals, human evals, pass rates, and quality scores
Evaluation suites
EvalGate traces screen with recent traces and a span timeline for agent and tool calls
Trace timeline
EvalGate workflow DAG showing multi-agent workflow nodes and handoff paths
Workflow DAG

Screenshots pulled from the private EvalGate org repo and verified against local code-rendered demo routes.

failure-to-gate loop
4-step
SDK surfaces
TS + Python
judge aggregation
5 modes
PII, retention, org policy
Controls
Shipped enterprise fintechTrust anchor

Chase Rewards Platform

Regulated rewards, enrollment, and partner platform work

PressureCustomers needed a faster way to enroll in benefits and understand rewards value while internal teams needed credible releases across mobile, partner, regulatory, and production-quality constraints.

My roleOwned product execution across enrollment, rewards, partner integration, mobile experience, release quality, and cross-functional tradeoffs with engineering, design, business, and compliance partners.

ProofDelivered two separate Chase workstreams: event-driven partner-benefit enrollment modernization, and native gift-card checkout migration measured through an A/B comparison.

  • regulated fintech
  • partner integration
  • mobile enrollment
  • customer activation
  • release discipline
  • production credibility

Metrics are grouped by workstream: enrollment modernization used operational before/after metrics; gift-card checkout used an A/B comparison against the existing experience.

Enrollment modernizationlower enrollment time
79%Cassandra-backed batch enrollment vs. Kafka event processing; web/mobile partner benefits.
Enrollment modernizationreduction in production errors
95%Legacy enrollment service vs. decoupled event-driven service; enrollment platform integrations.
Gift-card checkout migrationactivation lift
8%A/B: existing checkout vs. migrated native checkout; iOS/Android gift cards.
Gift-card checkout migrationrevenue lift
10%A/B: existing checkout vs. migrated native checkout; iOS/Android gift cards.
Featured systems

Smaller proof of range: context, support, platform controls, and growth execution.

These do not compete with EvalGate or Chase. They show breadth in AI workflows, support quality, internal platforms, and full-stack execution.

DreamFi ProductOS console showing topic rooms, source-grounded context, and operating metrics
AI workflow systemInternal prototypeLive demo

DreamFi ProductOS

Shared Banking Context and Insight Layer

A Glean-like banking intelligence layer for shared context, source-grounded answers, decision rooms, and reviewable product artifacts.

User problem: Banking teams lose decision quality when customer signals, product context, connector data, and generated artifacts are scattered across tools.

Built by Paul: Designed the context model, connector-backed question flow, evidence inspection, topic rooms, and eval signals for handoff-ready product work.

Target hypothesis: Faster product reviews, better source attribution, and fewer unsupported AI-generated decisions.

Proof signal: Source-grounded context, review states, eval signals, and reusable artifacts.

  • RAG
  • Connectors
  • Review states
  • Source attribution
Interface preview for a support intelligence dashboard with governance and cost signals
Support AIPersonal buildRepo evidence

Support101

LLM Support Intelligence Platform

A support quality system that keeps answer confidence, retrieval quality, escalation readiness, and cost visible to operators.

User problem: Support leaders need faster AI-assisted triage without hiding low-confidence answers, weak retrieval, or expensive model behavior.

Built by Paul: Built the support intelligence architecture, RAG workflow, governance controls, and operator-facing console.

Target hypothesis: Lower handle time, fewer unsupported AI responses, and clearer human escalation paths.

Proof signal: Governed RAG workflow, confidence surfaces, and escalation-ready review states.

  • FastAPI
  • Next.js
  • RAG
  • Governance
Interface preview for governed infrastructure workflows and policy checks
Platform controlsInfrastructure prototypeRepo evidence

Allstar Forge

Enterprise Data and Infrastructure Platform

A cloud-native platform pattern for analytics infrastructure, approvals, policy checks, observability, and operational handoffs.

User problem: Internal platform teams need reusable controls for data, approvals, and observability instead of brittle one-off scripts.

Built by Paul: Shaped the platform architecture, governance workflow, observability model, and approval-oriented UX.

Target hypothesis: Fewer production errors and cleaner handoffs through standardized policy and approval paths.

Proof signal: Policy checks, approval paths, observability hooks, and reusable infrastructure vocabulary.

  • Temporal
  • Policy checks
  • Approvals
  • Observability
Nudgr friction dashboard showing funnels, drop-off metrics, and impact scoring
Full-stack executionPersonal buildLive demo

Nudgr

Friction Analytics Platform

A full-stack product for detecting user friction, ranking drop-off patterns, and turning funnel pain into experiment priorities.

User problem: Growth teams can see drop-offs, but often lack the friction patterns, evidence, and priority order behind them.

Built by Paul: Owned the product concept, friction taxonomy, dashboard UX, Fastify services, and AI insight workflow.

Target hypothesis: Faster experiment prioritization and clearer activation work from ranked intervention opportunities.

Proof signal: Full-stack app, event model, ranked insights, and live deployment.

  • React
  • Fastify
  • Journey tracking
  • AI insights
Experiments / interview builds

Useful prototypes, clearly labeled so they do not dilute the flagship story.

These are concept builds and interview-style explorations. They preserve depth without implying a client relationship or shipped business result.

Interview buildConcept - unaffiliated

CVS Assistant

Healthcare AI Companion

A regulated-assistant concept exploring structured intake, user context, and safety-aware AI interaction.

Why it exists: Healthcare assistants need useful guidance while keeping intake, safety, and user context explicit.

Interview buildConcept - unaffiliated

Affirm

AI-Native Agentic Search

A fintech discovery concept that reframes search as a guided decision system for affordability, intent, and tradeoffs.

Why it exists: Fintech customers need explainable plan guidance, not a generic list of financing options.

ExperimentInfrastructure concept

Fin Observability

AI Fraud and Incident Platform

A fraud and incident response concept for signals, risk context, escalation evidence, and auditability.

Why it exists: Fraud and incident teams lose time when signals, risk context, and escalation evidence live in separate tools.

ExperimentLive prototype

Studio Pilot Vision

AI Portfolio Command Center

A portfolio operating concept for launch risk, dependencies, blocked work, and delivery momentum.

Why it exists: Product portfolios drift when leaders rely on status updates instead of risk, dependency, and momentum signals.

Additional build archive
  • AnomalyInfrastructure concept

    A monitoring concept for financial systems using ML models to detect irregular patterns and operational risk.

  • Trend WhisperPersonal build

    A retail trend analytics concept for signal collection, predictive forecasting, recommendations, and planning.

  • Glean Insight HubConcept - unaffiliated

    An enterprise search analytics concept for engagement insights, content gaps, recommendations, and admin workflows.

  • DreamFi EventsInternal prototype

    A repeatable growth system for event planning, partner coordination, follow-up, analytics, and acquisition funnels.

  • DreamFi DesignInternal prototype

    A governed source of truth for product and brand assets with semantic search, enrichment, and review states.

How I work with teams

I build momentum by making risk, ownership, and handoffs easier to see.

My best work happens in the middle of product, engineering, design, risk, data, and operations: clarifying the decision, making tradeoffs visible, and helping teams ship with confidence instead of ambiguity.

I align people before artifacts

I make the decision, user problem, owner, constraints, evidence, and tradeoffs explicit before a team burns cycles polishing the wrong thing.

I translate across functions

I move between engineers, design, data, risk, ops, compliance, partners, and executives without losing the customer thread or release constraint.

I keep delivery visible

I like crisp async updates, decision logs, demo loops, rollout checklists, rollback paths, and launch risks written down early enough to act on.

I protect quality without freezing teams

I push for reviewed evidence, metrics, guardrails, and release confidence while still finding the smallest useful increment the team can ship.

What teammates should feel from working with me

Fewer mystery priorities, fewer vague asks, clearer rollback paths, better handoffs, and a PM who can stay close to the details without turning every decision into a meeting. I try to be direct, useful, and safe to trust when the work gets ambiguous.

  • Clear problem framing
  • Calm prioritization
  • Tradeoffs over theater
  • Tight feedback loops
  • Useful docs
  • Rollback paths
  • Shared ownership
  • Respect for craft
  • Launch discipline