Be wary of WhatsApp messages impersonating Jobline Resources's staff offering job opportunities. Those who encounter suspicious messages can contact Jobline at +65 6339 7198

Responsibilities

  • Design, build, optimise, and maintain batch and streaming data ingestion pipelines using platforms such as Databricks and Kafka, ensuring scalability, reliability, observability, and alignment with enterprise data architecture standards.
  • Perform data transformation and cleansing using PySpark or SQL based on business and technical requirements
  • Monitor and troubleshoot data workflows to ensure data quality and pipeline reliability
  • Provide technical guidance to engineers and delivery partners on data platform patterns, reusable components, code quality, deployment readiness, and production support practices.
  • Lead integration of data from diverse source systems including files, APIs, databases, and streaming platforms, working with source-system owners and consuming teams to define fit-for-purpose ingestion patterns and delivery timelines.
  • Help maintain metadata and pipeline documentation for transparency and traceability
  • Own production readiness for assigned data and AI platform components, including observability, incident triage, root-cause analysis, release coordination, and continuous improvement of operational runbooks.
  • Participate in integrating pipelines with tools such as Microsoft Fabric, Databricks, Delta Lake, and other platform components
  • Build and maintain knowledge base and RAG solution on variety of hosting platforms
  • Implement and operate knowledge base storage, lifecycle management and embedding/vectorization
  • Contribute to automation efforts using version control and CI/CD workflows
  • Apply data governance, security, access control, and operational risk policies during solution design and implementation, ensuring pipelines and knowledge platforms meet enterprise compliance requirements.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field
  • 5–8 years of experience in data engineering, data platform engineering, or cloud-scale analytics solution delivery, with demonstrated ownership of production pipelines and platform components.
  • Proven ability to independently design, build, optimise, and operate production-grade batch or streaming data pipelines, including orchestration, observability, error handling, performance tuning, and operational support.
  • Hands-on experience with Python and SQL for data transformation and validation
  • Familiarity with Apache Spark (especially PySpark) and large-scale data processing concepts
  • Experience with implementing knowledge base and RAG solutions for agentic AI use cases
  • Self-starter with strong problem-solving skills and a keen attention to detail
  • Able to work independently and lead technical discussions with engineers, architects, product owners, source-system teams, and business stakeholders to translate requirements into secure and maintainable platform solutions.
  • Strong documentation and communication skills
  • Strong understanding of enterprise data architecture, cloud security, access control, CI/CD, release management, and production operations for data and AI platform solutions.