Data Engineer
Job Description
<div class="content-intro"><div> <div> </div> <div>At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life better for people. Our 13,000+ people are united in a shared purpose to pursue the limitless potential of technology and ingenuity for clients across defense, national security, public safety, civilian, and military health organizations. </div> <div> </div> <div>Join Accenture Federal Services, a technology company within global Accenture. Recognized as a Glassdoor Top 100 Best Place to Work, we offer a collaborative and caring community where you feel like you belong and are empowered to grow, learn and thrive through hands-on experience, certifications, industry training and more. </div> <div> </div> <div>Join us to drive positive, lasting change that moves missions and the government forward!</div> <div> </div> </div></div><p>We are looking for a <strong>Data Engineer</strong> with strong hands‑on experience designing, developing, and managing large‑scale data workflows across structured and unstructured datasets. This role focuses heavily on building reliable RAG (Retrieval-Augmented Generation) pipelines, orchestrating ETL/ELT processes, and deploying scalable data systems in AWS.</p> <p><strong>Responsibilities </strong></p> <ul> <li>Design, build, and maintain RAG pipelines, including document ingestion, indexing, embedding workflows, and model retrieval flows.</li> <li>Develop and manage structured and unstructured data pipelines supporting analytics, ML, and application workloads. Build and optimize ETL/ELT pipelines in AWS using services such as S3, Lambda, Step Functions, EMR, Glue, ECS/EKS, and IAM best practices. Implement and operate NiFi flows for high‑throughput, low‑latency data ingestion and transformation.</li> <li>Develop, orchestrate, and schedule workflows using Prefect, ensuring reliability, observability, and proper error handling. Implement indexing, search, and retrieval patterns using ElasticSearch, including schema design, cluster management, and query optimization.</li> <li>Collaborate closely with architecture, ML, and application teams to support scalable data solutions.</li> <li>Ensure data quality, lineage, governance, and security across all pipelines.</li> <li>Monitor system performance and troubleshoot issues across distributed data systems.</li> </ul> <p><strong>Qualifications</strong></p> <ul> <li>Solid understanding of ETL/ELT processes and data modeling best practices.</li> <li>Hands‑on experience implementing workflows in Prefect (Prefect 2.0 preferred). In‑depth knowledge of ElasticSearch indexing, cluster management, and search optimization.</li> <li>Proficiency in Python and familiarity with common data libraries (Pandas, PySpark, requests, etc.).</li> </ul> <p><strong>Preferred Qualifications </strong></p> <ul> <li>Strong experience with Apache NiFi for data flow management and real‑time ingestion.</li> <li>Experience building or maintaining RAG pipelines (e.g., vector databases, embeddings, document chunking strategies, retrieval optimization).</li> <li>Proven ability to manage structured and unstructured data pipelines at scale.</li> <li>Experience with AWS cloud services for data engineering.</li> <li>Strong version control and CI/CD experience.</li> <li>Experience with vector databases (OpenSearch, Pinecone, Weaviate, etc.)</li> <li>Familiarity with containerized workflows (Docker, Kubernetes)</li> <li>Experience supporting LLM or generative A