Principal Data Engineer
Job Description
<p><span style="color: #1f3864;"><strong>Overview</strong></span></p> <p style="color: !important;"><span style="color: #000000;">The Principal Data Engineer is a senior technical leader within AvidXchange's Data Engineering organization responsible for architecting, building, and scaling modern data platforms. In this role, you will drive the migration of legacy Azure SQL Server workloads to Databricks, design real-time streaming pipelines with Apache Kafka, enable AI and agentic capabilities such as Databricks Genie, and define the long-term data architecture strategy. You will partner closely with Software Engineering, Product, Architecture, DevOps, and Analytics teams to deliver secure, high-performing, and reliable data solutions at enterprise scale.</span></p> <p> </p> <p style="color: !important;"><span style="color: #1f3864;"><strong>What You'll Do</strong></span></p> <p style="color: !important;"><span style="color: #2e74b5;"><strong>Data Platform Architecture & Modernization</strong></span></p> <p>·<span style="font-size: 7pt;"> </span>Lead the design and implementation of scalable, cloud-native data architectures on Databricks (Delta Lake, Unity Catalog, Lakehouse patterns).</p> <p>·<span style="font-size: 7pt;"> </span>Own and execute the migration strategy from legacy Azure SQL Server to Databricks, including schema translation, ETL/ELT re-platforming, data validation, and cutover planning.</p> <p>·<span style="font-size: 7pt;"> </span>Define data modeling standards (medallion architecture, star/snowflake schemas) and ensure consistency across all pipelines and domains.</p> <p>·<span style="font-size: 7pt;"> </span>Evaluate and recommend tools, frameworks, and platforms to support long-term data strategy and organizational goals.</p> <p>·<span style="font-size: 7pt;"> </span>Collaborate with Solution and Enterprise Architects to review and approve new data architecture designs.</p> <p style="color: !important;"><span style="color: #2e74b5;"><strong>Streaming & Real-Time Data Engineering</strong></span></p> <p>·<span style="font-size: 7pt;"> </span>Architect and implement Kafka-based streaming pipelines for real-time data ingestion, transformation, and delivery.</p> <p>·<span style="font-size: 7pt;"> </span>Design event-driven architectures and streaming topologies using Kafka Streams, ksqlDB, or Spark Structured Streaming on Databricks.</p> <p>·<span style="font-size: 7pt;"> </span>Establish patterns for schema management (Confluent Schema Registry), consumer group strategy, offset management, and dead-letter queuing.</p> <p>·<span style="font-size: 7pt;"> </span>Ensure streaming pipelines meet SLA requirements for latency, throughput, and fault tolerance.</p> <p style="color: !important;"><span style="color: #2e74b5;"><strong>Optimization, Quality & Standards</strong></span></p> <p>·<span style="font-size: 7pt;"> </span>Debug and optimize Spark jobs, Delta Lake tables, and SQL workloads for performance, cost efficiency, and maintainability.</p> <p>·<span style="font-size: 7pt;"> </span>Lead code reviews focused on senior engineers to enforce s