AI/ML Data Engineer- Hybrid (US Citizens/ Green Cards- Local to DMV only)

SwingtechWashington, District of ColumbiaOn-siteFull-timeMid level, 2–5 yearsListed 8 hours ago

Apply now

About this role

company

About Swingtech

Swingtech delivers innovative Information Technology and Professional Support services to a diverse range of clients across the federal and intelligence communities. With over 15 years of trusted experience as a systems integrator, we apply agile methodologies and deep industry insight to help our customers achieve greater efficiency, compliance, and cost savings. At Swingtech, we’re committed to excellence and long-term success for our clients and our team.

role

Position Summary

The AI/ML Data Engineer develops and sustains the secure data pipelines, data products, retrieval foundations, and governance controls that enable DOL’s AI/ML solutions. The position supports structured, semi-structured, and unstructured data sources used for analytics, AI/ML development, document intelligence, RAG, and production AI applications.

Essential Duties

- Design, build, test, deploy, and maintain scalable data pipelines for batch, streaming, near-real-time, and event-driven workloads.
- Integrate approved agency data sources, APIs, file stores, document repositories, relational databases, data lakes, data warehouses, and authorized external sources.
- Develop ETL/ELT pipelines for data extraction, validation, transformation, normalization, enrichment, de-identification, metadata management, and loading.
- Implement document-ingestion pipelines that support OCR, parsing, classification, metadata extraction, PII detection/redaction, chunking, embeddings, vector indexing, and retrieval workflows.
- Create and maintain data models, schemas, data dictionaries, metadata structures, catalog records, and data-quality controls.
- Implement data lineage, source provenance, dataset versioning, retention, access controls, and auditability for training, validation, evaluation, and production datasets.
- Preserve the separation of training, validation, and final evaluation datasets through controlled access, versioning, and documented lifecycle processes.
- Develop and monitor data-quality measures, including completeness, accuracy, timeliness, duplication, validity, freshness, distribution drift, and labeling quality.
- Apply data minimization, masking, encryption, access controls, de-identification, and least-privilege safeguards to PII, CUI, and other protected DOL data.
- Collaborate with AI/ML Engineers to optimize retrieval quality, embeddings, vector stores, hybrid search, reranking, citation traceability, and knowledge-base refresh processes.
- Develop data-pipeline runbooks, technical documentation, source inventories, lineage artifacts, data-quality reports, and operational support procedures.
- Support security, privacy, ATO, Responsible AI, incident response, MLOps, monitoring, and release-readiness activities.

Required Qualifications

- Bachelor’s degree in computer science, data engineering, data science, information systems, software engineering, mathematics, or a related technical discipline.
- At least four years of experience in data engineering, database development, analytics engineering, ETL/ELT development, data-platform implementation, or related work.
- Strong SQL and Python development skills.
- Experience designing data pipelines and integrating APIs, databases, file systems, cloud storage, data warehouses, or data lakes.
- Experience with data modeling, metadata, data quality, data lineage, data transformation, monitoring, and operational support.
- Familiarity with AWS, Azure, Google Cloud, or equivalent cloud data services.
- Knowledge of secure data-handling practices, including access control, encryption, data masking, PII protection, and logging.
- Must be willing to work 3 days onsite at customer site in Washington, DC.

Preferred Qualifications

- Experience with AWS Glue, S3, Athena, Redshift, Lake Formation, Azure Data Factory, Azure Data Lake Storage, Databricks, Snowflake, BigQuery, or equivalent platforms.
- Experience with vector databases or vector-search capabilities, including OpenSearch, pgvector, Pinecone, Weaviate, Milvus, Chroma, FAISS, or similar tools.
- Experience with RAG, document intelligence, OCR, enterprise search, knowledge management, document classification, or content-ingestion pipelines.
- Familiarity with Federal data governance, FedRAMP, FISMA, NIST 800-53, NIST 800-171, CUI, Privacy Act, and records-management requirements.

Summary of Benefits

- 15 PTO days
- 11 paid holidays
- Medical Insurance with – 3 options (HSA with $600 Employer Contribution).
- Dental Insurance with no age limit orthodonture.
- Vision Insurance through EyeMed in and out of network coverage.
- Short Term and Long-Term Disability coverage with 100% premium support,
- Life insurance and AD&D with 100% premium support
- Supplemental Life Insurance
- Critical Care and Accident Insurance availability
- Pet Insurance through Nationwide
- Employee Assistance Program
- 401k with enrollment from day one. 4% deferral by company.
- $1500 Annual Training Budget
- $1500 Referral bonus
- Eligibility for annual merit and discretionary bonus
- Flexible work arrangements

Equal Opportunity Employer Minority/Female/Veterans/Disabled