Data Scientist

Rice UniversityUnited StatesOn-siteFull-timeJunior, 1–2 yearsListed 3 hours ago

Apply now

About this role

The Rice University Department of Computer Science and Department of Materials Science and NanoEngineering are seeking a Data Scientist to build and operate the data infrastructure for READINESS, a new $20 million NSF-funded project transforming four materials synthesis systems (CVD, thermal, PVD, and plasma CVD reactors for 2D materials, oxides, and diamond films) into an AI-driven, remotely accessible autonomous laboratory.

The READINESS facility will generate substantial and diverse data, including growth recipes, reactor time-series, user interactions, optical, SEM/TEM and AFM images, spectra, and continuous robot telemetry. The Data Scientist will be responsible for developing and maintaining the central data lake that captures, indexes, secures, and serves this information across the project.

The position combines data engineering and analysis, including deploying an S3-compatible object store, developing ingestion connectors for live instruments and robot controllers, and establishing SQL and vector-search capabilities under Rice single sign-on. The role will also compute embeddings over images and spectra, develop retrieval and similarity-search services, and build AI agents capable of answering researchers' questions directly from the data.

Reporting to the faculty leads of the READINESS Data Infrastructure working group, the Data Scientist will collaborate closely with the synthesis platform teams, robotics and AI/digital-twin working groups, Rice IT and Information Security, and graduate and undergraduate students involved in the project.

This is a hands-on computer science and data engineering position that requires substantial ownership of the project's central data infrastructure. A background in materials science, chemistry, or laboratory science is not required ; domain-specific knowledge can be developed through collaboration with the project's scientific experts. Working within a small team, the Data Scientist will develop the foundational systems through which data generated across the READINESS facility is captured, organized, and made accessible.

Special Instructions to Applicants:

Submit the following with your Rice application:

(1) CV

(2) a short statement explaining how your experience aligns with the position requirements; and

(3) three references including contact information.

Ideal Candidate Statement:

The ideal candidate is a skilled software or data engineer who enjoys collaborating across disciplines and brings experience designing data models and developing reliable systems to capture, organize, and serve complex data. Equally comfortable deploying storage infrastructure, troubleshooting connections to physical instruments, training embedding models, and integrating LLM agents with databases, this individual has experience developing and deploying data systems that others rely on. An appreciation for the challenges and opportunities presented by diverse, real-world data generated by scientific equipment is essential, along with careful attention to system reliability, documentation, and data integrity.

Self-directed and pragmatic, they can translate a project goal into a working technical solution, selecting appropriate tools and collaborating with the researchers who operate the instruments to ensure data is captured accurately and consistently with the necessary metadata. They communicate effectively with colleagues from different technical backgrounds and understand the importance of security, provenance, and access control in a multi-institutional research environment.

Workplace Requirements:

This position is fully remote, permitting all tasks to be completed from any location within the United States. Working hours will remain central standard time. Per Rice policy 440, work arrangements may be subject to change.

Hiring Range: Up to $105,000 annually.

This is a one-year term limited, benefits-eligible position funded by a grant, soft and/or restricted funds, and may be renewed based on continued funding availability, research needs, and performance.

Minimum Requirements:

- Bachelor’s degree in computer science, data science, engineering, or a related quantitative field.
- Three or more (3+) years of related professional experience in data engineering, data science, software engineering, or research computing.

Skills:

- Strong programming skills and ability to write production-quality, tested, documented code.
- Proficiency with SQL and relational data modeling, including scientific or operational database schemas.
- Experience with object storage (S3, Ceph, or MinIO), columnar formats (Parquet), and distributed query engines (Trino, Presto, Spark, or similar).
- Working knowledge of Linux administration, containers (Docker/Kubernetes), and at least one major cloud platform (AWS or Azure).
- Familiarity with PyTorch, embedding models, vector databases, and retrieval-augmented or agentic LLM applications.
- Understanding of authentication and access controls (SSO/SAML/OIDC, role-based access, audit logging) and secure research-data handling.
- Ability to connect software to physical instruments and devices using network or file-based protocols; ROS/rosbag experience is a plus.
- Ability to learn an unfamiliar scientific domain, translate collaborators’ requirements into working systems, manage competing requests, and communicate clearly.

Preferences:

- Five or more years of professional experience building and operating data systems.
- Experience building data pipelines for scientific instruments, laboratory automation, manufacturing, or IoT/telemetry.
- Experience operating Ceph or comparable software-defined storage, or managing cloud object-storage tenancies at 100 TB+ scale.
- Experience with GPU-accelerated similarity search (FAISS, Milvus, Qdrant) and serving open-weight LLMs.
- Experience in academic research or national laboratories, including working with institutional IT and security offices.
- Experience monitoring production data systems and implementing backup and recovery.

Essential Functions:

- Designs, deploys, and operates the READINESS central data lake, including the S3-compatible object store (self-hosted Ceph or cloud), bucket layout, versioning, and access policies.
- Develops and maintains ingestion connectors that capture data unattended from synthesis tools, characterization instruments, and robotic systems, and land it in the store with structured, de-identified metadata and provenance links.
- Designs and maintains the project's metadata and provenance schema and the Python client library that all project teams use to read and write data.
- Deploys and administers the SQL query layer (Trino and Hive Metastore over Parquet) and programmatic APIs for researchers, the user portal, and digital-twin and AI teams.
- Builds embedding pipelines and GPU-resident vector indexes over images, spectra, and logs, and exposes similarity-search and retrieval-augmented-generation services.
- Develops and demonstrates AI agents that answer questions about experiments by retrieving from the data lake, in collaboration with the AI & Digital Twins working group.
- Implements and maintains security controls, including Rice single sign-on with MFA, tiered role-based access, encryption, and immutable audit logging, and works with Rice Information Security and Research Security on design reviews and compliance.
- Monitors system health, performs integrity checks and backups, and documents and tests recovery procedures.
- Gathers requirements from synthesis platform, characterization, and robotics teams; attends platform team and working-group meetings; and sets node-wide standards for data formats, metadata, and access.
- Writes documentation, user guides, and onboarding materials, and trains researchers and partner-site collaborators to use the data infrastructure.
- Mentors graduate and undergraduate students working on data infrastructure projects.
- Performs all other related duties as assigned

Rice University HR | Benefits: https://knowledgecafe.rice.edu/benefits
Rice Mission and Values : Mission and Values | Rice University

Rice University is committed to ensuring Equal Employment Opportunity and welcoming the fullness of diversity into our candidate pools. Rice considers qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national or ethnic origin, genetic information, disability, or protected veteran status. Rice also provides reasonable accommodations to qualified persons with disabilities. If an applicant requires a reasonable accommodation for any part of the application or hiring process, please get in touch with Rice University’s Human Resources Office via email at [email protected] for support.

If you have any additional questions, please email us at [email protected] . Thank you for your interest in employment with Rice University.