Skip to content
All jobs

Full Time: Lead Data Engineer

Digital Minds Global Technologies Inc.

Job
29370
Posted
Location
Cary, NC
Work type
Full Time
Tax terms
W2, Yearly
Experience
Experience open
Openings
1 opening

Opens your email app with a message to the employer, its subject naming this job. Attach your resume and send it from your own email.

Skills

  • Kafka
  • SQL
  • Python
  • PySpark
  • Spark
  • Scala
  • Databricks
  • Delta Lake
  • Azure
  • Terraform
  • CI/CD

About the job

Lead Data Engineer (Hands-On)

Location: Cary, NC | On-site / Hybrid

Experience: 12-18 Years

Employment: Full-Time

Salary: $140K - $145K per annum + Benefits

About the Opportunity

We are looking for a hands-on Lead Data Engineer to support a centralized, AI-first enterprise Data Hub for a global insurance and financial services organization.

The platform is built on Azure Databricks, ingests 150+ inbound data feeds, and distributes data to 35+ downstream systems using a Bronze / Silver / Gold medallion architecture.

AI is embedded across ingestion, canonical mapping, data quality, reconciliation and business-user access.

This is a senior technical leadership role requiring hands-on coding. The successful candidate will own the end-to-end technical design of the data and AI layers, build reference implementations, review production code and deliver production-grade Python, Scala and PySpark solutions.

Candidates must have written or reviewed production code within the past year.

Key Responsibilities

• Design and implement enterprise lakehouse architecture using Bronze, Silver and Gold data layers

• Design ADLS Gen2 zones, Delta Lake tables, partitioning, schema evolution and retention strategies

• Build metadata-driven and parameterized ingestion frameworks for batch, database extracts, CDC and streaming data

• Develop production-grade Python, Scala and PySpark solutions

• Work with Azure Event Hubs, Kafka and Spark Structured Streaming

• Establish coding, testing and PR review standards

• Troubleshoot production incidents and optimize Spark workloads and cluster performance

• Implement CI/CD for Databricks and ADF using Azure DevOps, Terraform and Databricks Asset Bundles

• Build AI-augmented ingestion and source-to-canonical mapping solutions

• Implement AI-driven data quality, anomaly detection and automated reconciliation

• Develop synthetic, privacy-preserving test data solutions

• Contribute to semantic-layer, knowledge-graph and GPT-powered conversational data access capabilities

• Implement text-to-SQL and semantic retrieval with row- and column-level security

• Establish governance using Unity Catalog, lineage, access controls and PII standards

• Participate in architecture and AI governance forums

• Mentor engineering teams and provide technical leadership

Must-Have Skills & Experience

• 12-18 years of experience in data engineering / data platform delivery

• Expert-level Python, Scala and PySpark

• Strong SQL and data modelling skills

• Deep expertise in Databricks, Delta Lake and Unity Catalog

• Experience with Databricks Jobs & Workflows, cluster management and performance tuning

• Strong Azure experience including:

o ADLS Gen2

o Azure Data Factory

o Azure Event Hubs

o Azure security and governance

• Proven experience designing and delivering enterprise-scale medallion / lakehouse architectures

• 3+ years of production experience designing and implementing LLM-based systems

• Strong knowledge of RAG, agentic/tool-calling workflows, embeddings, vector/hybrid retrieval and prompt engineering

• Hands-on experience with LangChain, LlamaIndex or LangGraph

• Experience with Azure OpenAI, OpenAI or Databricks Model Serving

• Experience implementing evaluation frameworks including golden datasets, regression testing, accuracy measurement and hallucination tracking

• Experience with metadata-driven frameworks, schema inference, profiling, lineage and catalogs

• Strong experience with CI/CD and IaC using Azure DevOps, Terraform and Databricks Asset Bundles

• Knowledge of Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints and PII handling

• Strong technical communication and ability to present architecture to both technical and business stakeholders

Strongly Preferred

Similar jobs

See all jobs