Lead Data Engineer (Hands-On)
Cephas Consultancy Services Private Limited
United States, North Carolina, Cary
Employment: Full-time
Work arrangement: On-site or hybrid
Salary: Not specified
Eligibility: U.S. citizens only
About the Engagement
Build a centralized, AI-first enterprise data hub for a global insurance and financial services client on Azure Databricks. The platform ingests 150+ inbound data feeds, distributes data to 35+ downstream systems, and uses a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business-user access from day one.
This is a senior, hands-on leadership role. You will own the end-to-end technical design of the data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code every week. Candidates must have written or reviewed production code within the past year.
What the Role Owns
- Data platform architecture and engineering: Lakehouse architecture (Bronze / Silver / Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, and retention).
- Ingestion: Metadata-driven, parameterized frameworks for batch files, database extracts, CDC feeds, and streaming using Azure Event Hubs / Kafka and Spark Structured Streaming.
- Engineering practices: Canonical PySpark and Scala Spark jobs, coding and testing standards, PR reviews, production incident debugging, Spark cluster tuning, and cost guardrails.
- CI/CD and observability: Databricks and ADF in Azure DevOps using Databricks Asset Bundles and Terraform; Azure Monitor and Log Analytics.
- AI-augmented ingestion and canonical mapping: Auto-generated bridge documents, DML, and canonical table definitions; AI-assisted source-to-canonical mapping with a human review gate.
- AI-driven data quality: Anomaly detection (data drift, schema drift, volume shifts, and reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data.
- Semantic access: Semantic layer and knowledge graph, plus a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security.
- Governance and leadership: Unity Catalog (lineage, access control, and PII standards), Architecture Review Boards and AI governance forums, engineer mentoring, and documentation.
Must-Have Skills and Experience
- Expert-level Python, Scala, and PySpark, including production-ready, modular, well-tested solutions; Spark workload troubleshooting; and optimization of large-scale batch and streaming pipelines using Delta Lake.
- Strong SQL and data modelling (dimensional and normalized), schema design, and data contracts.
- Databricks expertise, including Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, and Model Serving.
- Experience with the Azure data stack: ADLS Gen2 (zone design, ACLs, and lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks), and Azure Event Hubs.
- At least 3 years designing and shipping LLM-based systems in production, including RAG pipelines, agentic / tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering.
- Evaluation experience with golden datasets, regression suites, accuracy and hallucination tracking, and human-in-the-loop feedback.
- Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
- Experience with metadata-driven frameworks, including schema inference, data profiling, lineage, and catalogs.
- 12–18 years of total experience in data engineering / data platform delivery.
- Proven enterprise-scale delivery of a medallion / lakehouse architecture.
- Azure security and governance experience, including Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, and PII handling.
- CI/CD and infrastructure-as-code experience with Azure DevOps, Terraform, Databricks Asset Bundles, and automated data-pipeline testing.
- Clear technical writing skills and the ability to present and defend designs to engineers and non-technical stakeholders.
Strongly Preferred
- Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, or graph modelling over a lakehouse).
- Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale.
- ML-based anomaly detection on time-series or transactional financial data.
- Financial services or insurance experience (finance close, GL, subledger, reconciliation, or actuarial data).
- LLMOps / MLOps experience, including model and prompt versioning, cost governance, and observability.
- Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.
- Experience with dbt, Great Expectations or similar tools; Workday, Prism, or Accounting Center exposure.