Be wary of WhatsApp messages impersonating Jobline Resources's staff offering job opportunities. Those who encounter suspicious messages can contact Jobline at +65 6339 7198
Design, build, optimise, and maintain batch and streaming data ingestion pipelines using platforms such as Databricks and Kafka, ensuring scalability, reliability, observability, and alignment with enterprise data architecture standards.
Perform data transformation and cleansing using PySpark or SQL based on business and technical requirements
Monitor and troubleshoot data workflows to ensure data quality and pipeline reliability
Provide technical guidance to engineers and delivery partners on data platform patterns, reusable components, code quality, deployment readiness, and production support practices.
Lead integration of data from diverse source systems including files, APIs, databases, and streaming platforms, working with source-system owners and consuming teams to define fit-for-purpose ingestion patterns and delivery timelines.
Help maintain metadata and pipeline documentation for transparency and traceability
Own production readiness for assigned data and AI platform components, including observability, incident triage, root-cause analysis, release coordination, and continuous improvement of operational runbooks.
Participate in integrating pipelines with tools such as Microsoft Fabric, Databricks, Delta Lake, and other platform components
Build and maintain knowledge base and RAG solution on variety of hosting platforms
Implement and operate knowledge base storage, lifecycle management and embedding/vectorization
Contribute to automation efforts using version control and CI/CD workflows
Apply data governance, security, access control, and operational risk policies during solution design and implementation, ensuring pipelines and knowledge platforms meet enterprise compliance requirements.
Requirements
Bachelor’s degree in Computer Science, Engineering, or a related field
5–8 years of experience in data engineering, data platform engineering, or cloud-scale analytics solution delivery, with demonstrated ownership of production pipelines and platform components.
Proven ability to independently design, build, optimise, and operate production-grade batch or streaming data pipelines, including orchestration, observability, error handling, performance tuning, and operational support.
Hands-on experience with Python and SQL for data transformation and validation
Familiarity with Apache Spark (especially PySpark) and large-scale data processing concepts
Experience with implementing knowledge base and RAG solutions for agentic AI use cases
Self-starter with strong problem-solving skills and a keen attention to detail
Able to work independently and lead technical discussions with engineers, architects, product owners, source-system teams, and business stakeholders to translate requirements into secure and maintainable platform solutions.
Strong documentation and communication skills
Strong understanding of enterprise data architecture, cloud security, access control, CI/CD, release management, and production operations for data and AI platform solutions.