About this role
Expected areas of expertise:
- Web crawlers and connectors for continuously changing international sources.
- Fault-tolerant pipelines with reliable scheduling, recovery, replay, and backfill capabilities.
- Processing complex documents, structured files, images, audio, and video across languages.
- Stable schemas, source lineage, correction propagation.
- Operating secure, observable, scalable, and cost-efficient processing systems across cloud and restricted environments.
- Document conversion, OCR, layout analysis, and structural extraction across complex file formats.
- Multilingual audio, video, and document processing with reliable alignment and attribution.
- Scalable inference infrastructure for batch and real-time processing across CPU and GPU workloads.
- End-to-end provenance, versioning, correction propagation, and reproducible reprocessing.
- Secure handling of untrusted content with rigorous quality evaluation, observability, and cost controls.
KEY RESPONSIBILITIES
- Expand Proxi’s ingestion and document-processing platform from source acquisition through downstream publication.
- Construct crawling and connector infrastructure required to discover, acquire, and continuously maintain international data sources.
- Scale the distributed processing environment that supports document conversion, extraction, transcription, enrichment, replay, and backfills.
- Establish common data contracts that normalize heterogeneous and multilingual sources without discarding their jurisdictional context or provenance.
- Maintain pipeline reliability, security, observability, and cost efficiency while ensuring corrections propagate through every downstream system.
- Build and operate real-time audio intelligence pipelines for low-latency multilingual transcription, speaker attribution, timestamp alignment, and confidence-calibrated tone analysis across live and recorded media.
NICE TO HAVE
- Experience building international public-sector data pipelines across languages and jurisdictions.
- Production experience with real-time transcription, speaker diarization, and tone or prosody analysis.
- Experience operating large-scale web acquisition against complex and frequently changing sources.
- Familiarity with legal, legislative, regulatory, archival, or similarly difficult corpora.
- Experience deploying secure processing systems in GovCloud, air-gapped, or other restricted environments.