dar4datascience
Menu

Career

Experience

Senior Data Engineer · TeamStation AI

May 2025 – Present · Remote (Mexico City) — consultancy for international clients

Architected an enterprise medallion data lake from scratch on AWS serverless and shipped self-service, natural-language analytics for business stakeholders.

  • Architected a ground-up medallion data lake (bronze/silver/gold) with custom Python packages and AWS Glue, Step Functions, Lambda, and DMS streaming.
  • Replaced Spark-heavy jobs with a DuckDB + Lambda architecture, cutting compute costs ~40% and processing time ~60% versus the equivalent Glue cluster approach.
  • Developed REST API endpoints with FastAPI and contributed to the Django backend, owning features end to end from API design to production deployment.
  • Offloaded analytics from the main application database into a dedicated PostgreSQL data warehouse and integrated ThoughtSpot BI for natural-language querying (NLQ).
  • Built an AI-enhanced Quarto documentation system covering Step Functions, Lambdas, and custom packages to speed up team onboarding.
  • Implemented CI/CD with GitHub Actions and infrastructure as code with AWS SAM, CloudFormation, and Terraform.
  • Python
  • SQL
  • PostgreSQL
  • AWS Glue
  • AWS Step Functions
  • AWS Lambda
  • AWS DMS
  • DuckDB
  • FastAPI
  • Django
  • ThoughtSpot
  • GitHub Actions
  • AWS SAM
  • CloudFormation
  • Terraform
  • Quarto

Business Intelligence Engineer IV · Rackspace Technology

2024 – 2025 · Remote (Mexico City)

Built enterprise ETL automation on BigQuery, created the organization's first Power BI CI/CD pipeline, and led a cost-optimization initiative worth over USD 1,000,000 per year.

  • Built Python ETL pipelines executing SQL against BigQuery and connecting 10+ enterprise APIs (Azure Monitor SDK, LeanIX OData, ServiceNow, SharePoint, Power BI); the AskHR pipeline ran hourly with zero manual intervention.
  • Independently designed and shipped an end-to-end Power BI CI/CD pipeline (Python + GitHub Actions + Azure AD service principals) for 90+ dashboards across 70+ workspaces, cutting deployment time from 60 minutes to 2 minutes.
  • Led BigQuery SQL performance tuning and decommissioning of stale assets, delivering estimated annual savings exceeding USD 1,000,000 and freeing 1,500+ engineering hours per year.
  • Built LLM-based applications for HR and BI use cases with OpenAI, Gemini, and agentic AI frameworks.
  • Worked with Databricks, Delta Lake, Unity Catalog, and GCP Dataproc/Dataflow for large-scale processing.
  • Python
  • SQL
  • BigQuery
  • Dataform
  • GitHub Actions
  • Power BI
  • Azure AD
  • Databricks
  • Delta Lake
  • PySpark
  • GCP Dataproc
  • GCP Dataflow
  • OpenAI API
  • Gemini API
  • Tableau
  • Jira

Senior BI Engineer · Baz Super App

2023 – 2024 · Mexico City

Led cross-cloud financial data streaming and automated marketing analytics on GCP.

  • Led sensitive financial data replication from AWS to BigQuery through a Kafka streaming pipeline, coordinating security, cloud, and data teams for compliant real-time availability.
  • Built and automated a Python ETL pipeline ingesting CleverTap marketing events into BigQuery, scheduled with GitHub Actions cron, enabling self-serve campaign reporting.
  • Implemented data validation and monitoring across cloud environments; worked with Databricks, Dataproc, Dataflow/Apache Beam, Pub/Sub, Cloud Functions, and Looker.
  • Python
  • SQL
  • BigQuery
  • Kafka
  • AWS
  • GCP
  • Cloud Composer
  • Pub/Sub
  • Apache Beam
  • Databricks
  • Looker
  • CleverTap
  • GitHub Actions

Data Engineer / BI Lead · DiDi Food

2021 – 2023 · Mexico City

Optimized Spark SQL workloads, automated analyst reporting, and led a SQL competency program for the team.

  • Tuned Spark SQL jobs via Python profiling and choke-point debugging, achieving a 40% reduction in job execution time.
  • Parameterized and automated SQL queries previously run manually by business users, saving ~100+ analyst hours per quarter.
  • Maintained data governance across customer-experience analytics pipelines spanning SQL Server and AWS sources (Athena, Glue, S3).
  • Designed and led a query competency program that raised the team's internal performance score and unlocked additional compute allocation.
  • Python
  • SQL
  • Spark
  • Presto
  • Hive
  • AWS Athena
  • AWS Glue
  • AWS S3
  • Airflow
  • dbt
  • Tableau
  • SQL Server
  • GitLab
  • Linux

Data Engineer · DGTIC UNAM

2019 – 2021 · Mexico City

Built the organization's first data warehouse for online lectures and pioneered Spark adoption for MOOC analytics.

  • Pioneered Spark adoption and built the first data warehouse for online lectures during the pandemic.
  • Collaborated with IT on database architecture and queried MOOC data from PostgreSQL for Big Data analytics.
  • Python
  • SQL
  • R
  • Spark
  • PostgreSQL
  • Docker
  • Linux