About this role
This is a remote position.
Project Description
- Support secure, reliable, and compliant operation of cloud-hosted research applications, agentic AI systems, scientific SaaS platforms, and production services.
- Work across GCP, AWS, cloud containers, AI/ML models, MCP servers, APIs, scientific databases, and third-party SaaS platforms .
- Support scientists, developers, IT, cybersecurity, networking, data teams, architecture, and vendors in transitioning research prototypes into secure, repeatable, production-ready services.
Key Responsibilities
- Administer GCP/AWS cloud projects, IAM, service accounts, networking, storage, compute, quotas, and SaaS environments.
- Operate Cloud Run, Docker containers, registries, DNS, TLS, load balancing, and Google Cloud IAP .
- Manage identities, RBAC, SSO, OAuth, service accounts, machine-to-machine authentication, secrets, API keys, tokens, and certificates.
- Maintain approved AI agents, MCP servers, tools, integrations, data sources, permissions, and credentials .
- Maintain Dockerfiles, container images, Python/application dependencies, registries, runtime configurations, and rollback procedures.
- Monitor logs, metrics, traces, alerts, dashboards, health checks, API limits, model/tool failures, token usage, cloud costs, and service availability.
- Maintain CI/CD pipelines and infrastructure-as-code , including source control, security scanning, testing, approvals, and deployment traceability.
- Perform incident triage, escalation, root-cause analysis, backup/recovery, disaster recovery, and service continuity.
- Manage incidents, service requests, problems, and changes using ServiceNow and Jira .
- Support vulnerability remediation, patching, security investigations, audits, threat modeling, and risk assessments.
- Identify shadow AI services, unmanaged integrations, unapproved MCP servers, and overprivileged identities.
- Create operational runbooks, technical documentation, and lifecycle/ownership documentation.
- Help transition scientist-managed prototypes into secure, supportable production deployments.
Requirements
Mandatory Requirements
- Bachelor’s degree in Computer Science, Information Systems, Engineering, Technology, or related field.
- 3+ years of experience in systems/cloud administration.
- Production administration experience with GCP, AWS, Azure, or comparable cloud platforms .
- Hands-on experience with:
- Linux
- Docker/containers
- Python environments
- APIs
- Networking, DNS, and TLS
- IAM, SSO, OAuth, service accounts
- Secrets management
- Serverless/container orchestration such as Cloud Run or Kubernetes
- Knowledge of:
- CI/CD
- Infrastructure as Code
- Monitoring and observability
- Incident management
- Backup/recovery
- Patching
- Vulnerability remediation
- Experience with ServiceNow and Jira .
- Strong troubleshooting, documentation, communication, and cross-functional collaboration skills.
Preferred Skills
- GCP IAP, Cloud Run, Secret Manager, Artifact Registry, Cloud Monitoring .
- Vertex AI, Gemini, or other AI/ML platform administration .
- Experience supporting GCP + AWS hybrid/multi-cloud environments .
- Agentic AI, AI orchestration, MCP servers , AI APIs, or model operations.
- Scientific computing, bioinformatics, laboratory automation, multiomics, research data platforms, or scientific SaaS.
- Container scanning, SBOM, software supply-chain security, data integrity, and auditability.
- Cloud, systems engineering, ITSM, or cybersecurity certifications.