Staff/Senior Software Engineer, Document Intelligence & Data Infrastructure (Helios)

GovSignalsNew York City, New YorkOn-siteFull-timeSenior, 5–8 yearsListed 2 days ago

Apply now

About this role

Expected areas of expertise:

- Web crawlers and connectors for continuously changing international sources.

- Fault-tolerant pipelines with reliable scheduling, recovery, replay, and backfill capabilities.

- Processing complex documents, structured files, images, audio, and video across languages.

- Stable schemas, source lineage, correction propagation.

- Operating secure, observable, scalable, and cost-efficient processing systems across cloud and restricted environments.

- Document conversion, OCR, layout analysis, and structural extraction across complex file formats.

- Multilingual audio, video, and document processing with reliable alignment and attribution.

- Scalable inference infrastructure for batch and real-time processing across CPU and GPU workloads.

- End-to-end provenance, versioning, correction propagation, and reproducible reprocessing.

- Secure handling of untrusted content with rigorous quality evaluation, observability, and cost controls.

KEY RESPONSIBILITIES

- Expand Proxi’s ingestion and document-processing platform from source acquisition through downstream publication.

- Construct crawling and connector infrastructure required to discover, acquire, and continuously maintain international data sources.

- Scale the distributed processing environment that supports document conversion, extraction, transcription, enrichment, replay, and backfills.

- Establish common data contracts that normalize heterogeneous and multilingual sources without discarding their jurisdictional context or provenance.

- Maintain pipeline reliability, security, observability, and cost efficiency while ensuring corrections propagate through every downstream system.

- Build and operate real-time audio intelligence pipelines for low-latency multilingual transcription, speaker attribution, timestamp alignment, and confidence-calibrated tone analysis across live and recorded media.

NICE TO HAVE

- Experience building international public-sector data pipelines across languages and jurisdictions.

- Production experience with real-time transcription, speaker diarization, and tone or prosody analysis.

- Experience operating large-scale web acquisition against complex and frequently changing sources.

- Familiarity with legal, legislative, regulatory, archival, or similarly difficult corpora.

- Experience deploying secure processing systems in GovCloud, air-gapped, or other restricted environments.