Data Architect (Managed Services)
Job Description
<p><strong>ROLE OVERVIEW</strong></p> <p>We are looking for a seasoned Data Architect with deep expertise in enterprise data warehousing, ETL/ELT pipeline development, and business intelligence. You will be responsible for the end-to-end design and implementation of our data architecture — from data ingestion and cleansing to dimensional modeling, warehouse construction, and BI dashboard delivery. You will also establish robust data governance frameworks including data catalogs, lineage tracking, and security audit systems. This role requires strong hands-on capability combined with architectural vision, and prior experience in complex industries such as finance, international trade, or manufacturing is highly valued.</p> <p><strong>KEY RESPONSIBILITIES</strong></p> <p><strong><em>Data Acquisition, Cleansing & Integration</em></strong></p> <ul> <li>Design and implement enterprise-wide data collection strategies across heterogeneous source systems (ETRM, ERP, CRM, external APIs, flat files, and databases).</li> <li>Develop robust data cleansing and standardization pipelines to ensure data accuracy, consistency, and completeness across all data sources.</li> <li>Build data integration frameworks to consolidate data from disparate systems into a unified enterprise data warehouse, handling schema evolution and data drift.</li> <li>Implement Change Data Capture (CDC) mechanisms and incremental data loading strategies to ensure near real-time data freshness.</li> </ul> <p><strong><em>Data Warehouse Architecture & Modeling</em></strong></p> <ul> <li>Lead the design and construction of the enterprise data warehouse (EDW) using industry-proven methodologies (Inmon, Kimball, or Data Vault).</li> <li>Apply dimensional modeling expertise to design star schemas, snowflake schemas, and constellation models optimized for analytical query performance.</li> <li>architect data marts for specific business domains (finance, sales, supply chain, operations) with clear separation of concerns.</li> <li>Select and implement appropriate data warehouse technologies based on workload characteristics: batch analytics (Hive, Spark SQL), real-time analytics (ClickHouse, Apache Doris, StarRocks), or MPP databases (Greenplum, Vertica).</li> <li>Design layered data architectures (ODS → DWD → DWS → ADS / Bronze → Silver → Gold) with clear data flow and transformation logic at each layer.</li> </ul> <p><strong><em>ETL/ELT Pipeline Development</em></strong></p> <ul> <li>Design, build, and maintain scalable ETL/ELT pipelines using Apache Spark, Flink, Hive, MapReduce, or proprietary ETL tools (Informatica, Talend, Kettle).</li> <li>Implement workflow orchestration using Apache Airflow, DolphinScheduler, Azkaban, or Oozie to schedule, monitor, and manage data pipelines.</li> <li>Establish data quality frameworks: automated validation rules, anomaly detection, data profiling, and quality scorecards to ensure pipeline reliability and data trustworthiness.</li> <li>Build comprehensive monitoring and alerting systems for pipeline health, data freshness, and SLA compliance.</li> <li>Optimize pipeline performance through partitioning, bucketing, indexing, and query tuning strategies.</li> </ul> <p><strong><em>Business Intelligence & Reporting</em></strong></p> <ul> <li>Develop enterprise BI reporting systems and interactive dashboards using tools such as Apache Superset, FineBI, Tableau, Power BI, or Quick BI.</li> <li>Design executive dashboards, operational reports, and self-service analytics interfaces tailored to business stakehol