Bitget
About us
Bitget is the world's largest Universal Exchange (UEX), serving over 125 million users and offering access to over 2M crypto tokens, 100+ tokenized stocks, ETFs, commodities, FX, and precious metals such as gold. The ecosystem is committed to helping users trade smarter with its AI agent, which co-pilots trade execution.
Bitget is driving crypto adoption through strategic partnerships with LALIGA and MotoGP™. Aligned with its global impact strategy, Bitget has joined hands with UNICEF to support blockchain education for 1.1 million people by 2027. Bitget currently leads in the tokenized TradFi market, providing the industry's lowest fees and highest liquidity across 150 regions worldwide.
What you'll do
1. Be responsible for the design, development, maintenance and continuous optimization of the company's offline data warehouse, supporting business scenarios such as transactions, market data, assets, users, risk control, compliance, growth and on-chain data.
2. Based on business requirements, conduct data modeling and build and maintain data warehouses at different levels such as ODS, DWD, DWS, ADS, etc., to accumulate common dimensions, detailed facts, summary indicators and data service models.
3. Be responsible for the development of offline ETL tasks, including data access, cleaning and transformation, detailed processing, indicator summarization, historical data backtracking, data filling and task dependency orchestration, etc.
4. Build a unified indicator system, data dictionary and indicator口径 management mechanism, promoting the unification of data standards across business teams and the reuse of data assets.
5. Be responsible for the performance optimization and cost management of offline computing tasks such as MaxCompute/Hive, including SQL optimization, partition design, Join optimization, data skew handling, table storage optimization and resource usage optimization.
6. Be responsible for the scheduling of offline tasks, SLA management, monitoring alerts and anomaly troubleshooting, ensuring that core data is produced on time, accurately and stably.
7. Build and improve the data quality system, including data integrity, accuracy, consistency, timeliness, volatility detection and upstream/downstream data verification, etc.
8. Participate in data governance work, including metadata management, data lineage, data permissions, data lifecycle, data security and historical data governance.
9. Collaborate with business, product, algorithm, R&D and risk control teams to complete requirements analysis, technical solution design, development delivery, data acceptance and subsequent iterative optimization.
10. Based on the needs of business and team technological evolution, participate in the exploration and construction of real-time data links, lakehouse integration or AI-assisted data research and development.
What you'll need
1. At least 5 years of experience in data development, data warehouse or big data development, with experience in offline data warehouse construction and maintenance in production environment.
2. Possess solid data warehouse modeling skills, familiar with dimension modeling, topic domain division, fact table and dimension table design, wide table design, summary table design, zipper table and snapshot table design, etc. modeling methods.
3. Familiar with the offline data warehouse hierarchical system, able to independently complete the design of hierarchical models such as ODS, DWD, DWS, ADS, etc. and the development of complex business data links.
4. Proficient in SQL, with the ability to write, optimize and troubleshoot complex SQL, familiar with Join optimization, window functions, partition trimming, data skew and big table processing optimization.
5. Skilled in using offline data warehouse and scheduling tools such as MaxCompute, DataWorks, Hive, etc., with experience in task development, scheduling orchestration, data replenishment and backfilling, task operation and cost optimization.
6. Familiar with big data ecosystem components such as Hadoop, Hive, Spark, Kafka, HBase, etc., and understand the basic principles of distributed storage, computing and resource scheduling.
7. Have experience in data quality, task monitoring, alarm handling, data governance, metadata or lineage management.
8. Familiar with at least one development language such as Python or Java, with good engineering capabilities and code standard awareness.
9. Have strong business understanding and data sensitivity, able to abstract complex business requirements into maintainable and scalable data models and indicator systems.
10. Have good communication and collaboration skills, sense of responsibility and problem-solving ability, able to independently locate and drive the resolution of complex data link problems.
[Extra points]
• Have data warehouse construction experience in Web3, blockchain, public chain, DeFi, cryptocurrency trading, finance, payment, risk control or compliance-related businesses.
• Have experience in TB/PB-level offline data processing, governance of large tables, task performance optimization or data cost governance.
• Familiar with Alibaba Cloud big data products, including MaxCompute, DataWorks, Hologres, DLF, EMR, OSS, etc.
• Have experience in real-time computing, OLAP or lakehouse-related experiences such as Flink, Kafka, Hologres, StarRocks, Paimon.
• Have experience in building data quality platforms, indicator platforms, metadata platforms, data development platforms, SQL review or automated operation tools.
• Have experience in large models, Agents, RAG, knowledge bases or AI-assisted R&D, and be able to use AI for SQL development, data quality checks, task troubleshooting or R&D efficiency improvement.
岗位职责:
1. 负责公司离线数据仓库的设计、开发、维护和持续优化,支撑交易、行情、资产、用户、风控、合规、增长及链上数据等业务场景。
2. 基于业务需求进行数据建模,建设并维护 ODS、DWD、DWS、ADS 等数仓分层,沉淀公共维度、明细事实、汇总指标及数据服务模型。
3. 负责离线 ETL 任务开发,包括数据接入、清洗转换、明细加工、指标汇总、历史数据回溯、补数及任务依赖编排等工作。
4. 建设统一的指标体系、数据字典及指标口径管理机制,推动跨业务团队的数据标准统一和数据资产复用。
5. 负责 MaxCompute/Hive 等离线计算任务的性能优化和成本治理,包括 SQL 优化、分区设计、Join 优化、数据倾斜处理、表存储优化及资源使用优化。
6. 负责离线任务调度、SLA 管理、监控告警及异常排查,保障核心数据按时、准确、稳定产出。
7. 建设和完善数据质量体系,包括数据完整性、准确性、一致性、及时性、波动检测及上下游数据核对等。
8. 参与数据治理工作,包括元数据管理、数据血缘、数据权限、数据生命周期、数据安全及历史数据治理等。
9. 与业务、产品、算法、研发及风控团队协作,完成需求分析、技术方案设计、开发交付、数据验收及后续迭代优化。
10. 结合业务和团队技术演进需要,参与实时数据链路、湖仓一体或 AI 辅助数据研发等相关探索与建设。
任职要求:
1. 5 年及以上数据开发、数据仓库或大数据开发相关经验,具备生产环境离线数仓建设和维护经验。
2. 具备扎实的数据仓库建模能力,熟悉维度建模、主题域划分、事实表和维表设计、宽表设计、汇总表设计、拉链表及快照表等建模方法。
3. 熟悉离线数仓分层体系,能够独立完成 ODS、DWD、DWS、ADS 等分层模型设计和复杂业务数据链路开发。
4. 精通 SQL,具备复杂 SQL 编写、调优和问题排查能力,熟悉 Join 优化、窗口函数、分区裁剪、数据倾斜及大表处理优化。
5. 熟练使用 MaxCompute、DataWorks、Hive 等离线数仓及调度工具,具备任务开发、调度编排、补数回刷、任务运维和成本优化经验。
6. 熟悉 Hadoop、Hive、Spark、Kafka、HBase 等大数据生态组件,理解分布式存储、计算和资源调度的基本原理。
7. 具备数据质量、任务监控、告警处理、数据治理、元数据或血缘管理相关经验。
8. 熟悉 Python 或 Java 至少一种开发语言,具备良好的工程化能力和代码规范意识。
9. 有较强的业务理解和数据敏感度,能够将复杂业务需求抽象为可维护、可扩展的数据模型和指标体系。
10. 具备良好的沟通协作能力、责任心和问题推进能力,能够独立定位并推动解决复杂数据链路问题。
【加分项】
• 有 Web3、区块链、公链、DeFi、加密货币交易、金融、支付、风控或合规等业务的数据仓库建设经验。
• 有 TB/PB 级离线数据处理、超大表治理、任务性能调优或数据成本治理经验。
• 熟悉阿里云大数据产品,包括 MaxCompute、DataWorks、Hologres、DLF、EMR、OSS 等。
• 具备 Flink、Kafka、Hologres、StarRocks、Paimon 等实时计算、OLAP 或湖仓相关经验。
• 有数据质量平台、指标平台、元数据平台、数据开发平台、SQL 审核或自动化运维工具建设经验。
• 有大模型、Agent、RAG、知识库或 AI 辅助研发经验,能够将 AI 用于 SQL 开发、数据质量检查、任务排障或研发提效。
Why Bitget?
If you are ambitious and believe that digital assets could be the next financial and technological revolution, please apply!
For more information regarding candidates' personal data processing, please refer to Bitget Candidate Privacy Notice.
Bitget