Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited

United States, North Carolina, Cary

Employment: Full-time

Work arrangement: On-site or hybrid

Salary: Not specified

Eligibility: U.S. citizens only

About the Engagement

Build a centralized, AI-first enterprise data hub for a global insurance and financial services client on Azure Databricks. The platform ingests 150+ inbound data feeds, distributes data to 35+ downstream systems, and uses a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business-user access from day one.

This is a senior, hands-on leadership role. You will own the end-to-end technical design of the data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code every week. Candidates must have written or reviewed production code within the past year.

What the Role Owns

  • Data platform architecture and engineering: Lakehouse architecture (Bronze / Silver / Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, and retention).
  • Ingestion: Metadata-driven, parameterized frameworks for batch files, database extracts, CDC feeds, and streaming using Azure Event Hubs / Kafka and Spark Structured Streaming.
  • Engineering practices: Canonical PySpark and Scala Spark jobs, coding and testing standards, PR reviews, production incident debugging, Spark cluster tuning, and cost guardrails.
  • CI/CD and observability: Databricks and ADF in Azure DevOps using Databricks Asset Bundles and Terraform; Azure Monitor and Log Analytics.
  • AI-augmented ingestion and canonical mapping: Auto-generated bridge documents, DML, and canonical table definitions; AI-assisted source-to-canonical mapping with a human review gate.
  • AI-driven data quality: Anomaly detection (data drift, schema drift, volume shifts, and reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data.
  • Semantic access: Semantic layer and knowledge graph, plus a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security.
  • Governance and leadership: Unity Catalog (lineage, access control, and PII standards), Architecture Review Boards and AI governance forums, engineer mentoring, and documentation.

Must-Have Skills and Experience

  • Expert-level Python, Scala, and PySpark, including production-ready, modular, well-tested solutions; Spark workload troubleshooting; and optimization of large-scale batch and streaming pipelines using Delta Lake.
  • Strong SQL and data modelling (dimensional and normalized), schema design, and data contracts.
  • Databricks expertise, including Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, and Model Serving.
  • Experience with the Azure data stack: ADLS Gen2 (zone design, ACLs, and lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks), and Azure Event Hubs.
  • At least 3 years designing and shipping LLM-based systems in production, including RAG pipelines, agentic / tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering.
  • Evaluation experience with golden datasets, regression suites, accuracy and hallucination tracking, and human-in-the-loop feedback.
  • Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
  • Experience with metadata-driven frameworks, including schema inference, data profiling, lineage, and catalogs.
  • 12–18 years of total experience in data engineering / data platform delivery.
  • Proven enterprise-scale delivery of a medallion / lakehouse architecture.
  • Azure security and governance experience, including Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, and PII handling.
  • CI/CD and infrastructure-as-code experience with Azure DevOps, Terraform, Databricks Asset Bundles, and automated data-pipeline testing.
  • Clear technical writing skills and the ability to present and defend designs to engineers and non-technical stakeholders.

Strongly Preferred

  • Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, or graph modelling over a lakehouse).
  • Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale.
  • ML-based anomaly detection on time-series or transactional financial data.
  • Financial services or insurance experience (finance close, GL, subledger, reconciliation, or actuarial data).
  • LLMOps / MLOps experience, including model and prompt versioning, cost governance, and observability.
  • Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.
  • Experience with dbt, Great Expectations or similar tools; Workday, Prism, or Accounting Center exposure.