Data Engineer I (R-19886)

Dun & BradstreetHyderabad, TelanganaOn-siteFull-timeNew grad, 0–1 yearsListed 2 weeks ago

Apply now

About this role

Shape the Future with Dun & Bradstreet
At Dun & Bradstreet, we believe data has the power to create a better tomorrow. As a global leader in business decisioning data and analytics, we help companies worldwide grow, manage risk, and innovate. Since 1841, businesses have trusted us to turn uncertainty into opportunity. We’re a diverse, global team that values creativity, collaboration, and bold ideas. Are you ready to make an impact and help shape what’s next? Join us! Explore opportunities at dnb.com/careers.

Key Responsibilities:

Coding & Development

- Write, review, test, and maintain code, including SQL and Python, to build data tools, automations, and ingestion and transformation workflows that support Research and Managed Services

- Automate manual processes and develop data tools to improve efficiency, accuracy, quality, and throughput

- Develop and promote coding standards and contribute to code reviews within the agile team

- Build and maintain web-scraping solutions, API integrations, and reusable data-processing components

- Support scalable ETL/ELT pipelines for structured and unstructured data

Source Evaluation & AI Enablement

- Identify prospective sources that can feed AI solutions and Research and Managed Services workflows

- Define and apply source-evaluation criteria covering relevance, authority, freshness, completeness, coverage, consistency, accessibility, legal or licensing constraints, privacy, security, and technical compatibility

- Perform source profiling, sample validation, proof-of-concept testing, and comparative assessments before recommending onboarding

- Document source decisions, metadata, lineage, ownership, limitations, refresh expectations, and approved use cases

- Implement and support AI-enabled workflows using LangChain or equivalent orchestration frameworks, large language models, embeddings, retrieval-augmented generation, vector databases, and prompt-engineering approaches where applicable

- Monitor source and AI-workflow performance and recommend remediation, replacement, or additional sources when quality or coverage falls below requirements

Operational Tasks

- Perform day-to-day operational activities supporting Research and Managed Services, including monitoring, exception handling, data maintenance, and issue resolution

- Perform database administration activities, including performance tuning and implementation of best practices

- Implement new data-maintenance processes and provide end-to-end process ownership

- Ensure data integrity by validating, reconciling, and regularly cleaning data

- Investigate and resolve production incidents, pipeline failures, data-quality issues, and operational exceptions

- Follow applicable data governance, security, and operational standards

Collaboration & Continuous Learning

- Evaluate and implement new technology solutions, and proactively learn and adopt new tools, platforms, and methodologies introduced by the organization

- Communicate with stakeholders and conduct knowledge-exchange sessions for technical and non-technical audiences

- Develop and maintain data documentation, including data dictionaries, source assessments, data-flow diagrams, data mappings, runbooks, and data lineage

- Collaborate with cross-functional teams across Data & Analytics, Technology, Research Services, Managed Services, Product, and Data Governance

- Additional duties as assigned.

Key Skills:

- Strong SQL and Python skills, with demonstrated ability to write and maintain code as a core part of daily work

- Experience with Playwright, Selenium, and other web-data collection techniques

- Experience developing and supporting data-ingestion, transformation, and ETL/ELT workflows

- Ability to collect and interpret data from multiple sources, including web scraping and GCS/S3, and formats including delimited files, XML, JSON, and PDF

- Working knowledge of data systems and databases used to maintain data pipelines

- Experience with Power BI, Tableau, or other dashboard tools

- Experience managing stakeholders and project plans

- Proficiency in Microsoft Office Suite

- Willingness and demonstrated ability to learn new technologies as they are introduced

- BigQuery experience and knowledge of AWS and/or GCP

- Hands-on experience implementing AI solutions using LangChain or an equivalent orchestration framework

- Exposure large language models, prompt engineering, retrieval-augmented generation, embeddings, vector databases, AI agents, or graph databases

- Knowledge of Data Operations methodologies, data management approaches, ServiceNow, and/or Jira

- Experience with NoSQL technologies, SQL Server administration, R programming, web technologies, and data mapping from multiple sources.

All Dun & Bradstreet job postings can be found at https://jobs.lever.co/dnb. Official communication from Dun & Bradstreet will come from an email address ending in @dnb.com.
 
Notice to Applicants: Please be advised that this job posting page is hosted and powered by Lever, a subsidiary of Employ Inc. Your use of this page is subject to Employ's Privacy Notice and Cookie Policy, which governs the processing of visitor data on this platform.