Product Manager, Stream & Batch Processing in Data Infrastructure

MetaLondon, EnglandOn-siteFull-timePrincipal, 12–15+ yearsListed 6 hours ago

Apply now

About this role

Meta’s products depend on data pipelines that process large-scale batch and real-time workloads. Today, those workloads span two ecosystems.

Our streaming stack includes XStream and a growing managed Apache Flink footprint. Our batch stack includes Spark and Presto over Hive and Iceberg tables, Velox as a shared execution layer, and an internal workflow authoring and scheduling platform.

These systems evolved separately. As a result, engineers encounter different authoring models, operational practices, guarantees, and migration paths depending on the workload they are building.

We are looking for a senior individual-contributor PM to set the product direction across this portfolio. The immediate focus is to define a coherent customer experience across batch and streaming, align the long-term platform strategy, and help teams move from legacy systems safely.

You will work closely with engineering leaders across multiple organizations. Success will require technical judgment, a strong understanding of internal customers, and the ability to turn a cross-organizational strategy into measurable adoption and operational improvements.

Responsibilities

Own the multi-year product strategy for stream and batch processing at Meta: the engine portfolio, the authoring surface, and the convergence path between streaming and batch execution.
Act as the voice of the internal customer: software engineers, data engineers and ML practitioners authoring pipelines.
Turn their real pain (state management, event-time windowing, backfill, schema evolution, debugging a pipeline at 3am) into abstractions that hold up under load.
Define what "managed" should mean internally, covering self-service authoring, autoscaling, upgrade and deprecation, and hold the platform to that bar.
Establish the semantics customers can rely on, and the honest cost of each guarantee: processing semantics (at-least-once, at-most-once, exactly-once), checkpointing, state management, and freshness.
Establish the service level objectives this platform commits to: throughput, freshness and lag, reliability, and compute efficiency.
Instrument whether the platform is measurably reducing operational load for the teams on it.
Own adoption and deprecation as a single problem.
Drive migration onto the converged path, build the onboarding surface that makes it the obvious choice, and retire legacy engines safely for customers with production systems, oncall rotations and revenue depending on the one you're turning off.
Drive the open source strategy: what we adopt, what we contribute upstream, and where internal systems remain justified.
Partner with the ML data and ranking teams whose realtime feature pipelines are the most demanding tenants on this stack, and with the warehouse and analytics teams on the batch side of the boundary.
Drive alignment across Ads, Reels, ML Infra and the warehouse orgs largely without direct authority.
Establish consensus on strategy and priorities with engineering leadership, and present to executive audiences.
Raise the bar of the PM function through mentorship of senior PMs across Data Infrastructure, and support recruiting and interviewing.

Qualifications

12+ years of experience in Product Management or equivalent relevant experience
Experience with or direct product ownership of at least one of Google Dataflow, Apache Beam, Flink, Kafka, Spark Structured Streaming, Iceberg or comparable systems
Demonstrated expertise in large-scale distributed data systems
Critical thinking and analytical leadership experience
Experience driving strategy and alignment across multiple organizations without direct authority
BA/BS in Computer Science or Information Systems Fluency in streaming semantics like event time versus processing time, windowing, watermarks, state backends, exactly-once delivery, deep enough to argue the tradeoffs
Hands on experience, for example, prior solutions architect or developer relations roles, is a strong plus
Experience owning autoscaling, multi-tenant capacity or compute efficiency as a product outcome on a fixed or constrained fleet
Experience with the ML side of streaming: realtime feature generation, freshness and coverage tradeoffs, training-serving consistency
A track record of successfully deprecating a system people depended on
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies