Skip to main content
find your future at ford.

Data Engineer

Job ID
69620
Category
Enterprise Technology
Location
India
Work Type
On-site

We are looking for someone with advanced Python programming skills who applies robust software engineering principles to data problems. You will collaborate closely with Data Scientists, ML Engineers, and Product Managers to build the scalable, automated pipelines required to train, deploy, and monitor machine learning models in production.

Key Responsibilities

  • GCP Pipeline Development: Design, build, and maintain highly scalable ETL/ELT data pipelines using Python and GCP-native data processing tools (e.g., Cloud Run, Cloud Functions).
  • AI/ML Infrastructure Support: Engineer feature stores, robust data feeds specifically optimized for machine learning training and inference. Work closely with ML Engineers to operationalize models using Vertex AI.
  • Data Integration & Ingestion: Write clean, modular Python code to ingest data from diverse sources (APIs, streaming platforms, on-prem databases) into BigQuery and Google Cloud Storage (GCS).
  • System Optimization: Optimize BigQuery architecture, partition/cluster tables, and tune complex SQL queries to ensure performance and cost-efficiency at a massive scale.
  • Software Engineering Best Practices: Champion best practices in Python development, including version control (Git), CI/CD pipelines (Cloud Build / GitHub Actions), code reviews, and comprehensive unit/integration testing.
  • Data Quality & Governance: Implement robust data quality checks, alerting, and monitoring to ensure the data feeding our AI models is accurate and reliable.

Required Qualifications

  • Degree: Bachelor’s or Master’s degree in Computer Science, Engineering, Mathematics, or a related technical field (or equivalent practical experience).
  • Experience: 4 to 6 years of professional experience in Data Engineering, Software Engineering, or a closely related field.
  • Advanced Python: Deep expertise in Python programming. You should be highly comfortable with:
    • Data processing and ML-adjacent libraries (e.g., PySpark, Pandas, NumPy).
    • API development
    • Writing efficient and production-grade code.
  • GCP Mastery: Proven, hands-on experience designing and operating data architectures on Google Cloud Platform. Must have strong experience with:
    • BigQuery (advanced SQL, architecture, and optimization).
    • Google Cloud Storage (GCS).
    • Compute/Serverless (Cloud Functions, Cloud Run).
  • AI/ML Acumen: Experience working alongside Data Science teams. A strong understanding of the ML lifecycle, feature engineering, and the data requirements for model training and deployment.
  • MLOps: Understanding of MLOps principles, model registry, and continuous training pipelines.

Preferred Qualifications

  • Vertex AI: Direct experience interacting with or deploying pipelines using Google Cloud's Vertex AI platform.
  • Streaming Technologies: Familiarity with real-time data processing using Google Cloud Pub/Sub and streaming Dataflow jobs.
  • Infrastructure as Code: Experience managing GCP resources using Terraform.
  • Containerization: Proficiency with Docker.
  • Built on one bold idea and the passion to define sustainable transportation for generations to come, Ford is a story about people with a vision that’s still being written.

    What We Do
  • Ford’s culture fuels the kind of momentum where ideas flow, progress is unstoppable, and our people keep redefining what it means to innovate.

    Our People and Culture
  • At Ford, your work matters, your life matters and we’re here to back the whole you—from growth to well-being—so you show up ready to realize your full potential.

    Your Benefits

Jobs For You.

Explore roles tailored to your interests, based on your preferences and experience.

Be the first to know about new jobs.

Sign Up Now