May 2025 – Present · Remote (Mexico City) — consultancy for international clients
Architected an enterprise medallion data lake from scratch on AWS serverless and shipped self-service, natural-language analytics for business stakeholders.
Architected a ground-up medallion data lake (bronze/silver/gold) with custom Python packages and AWS Glue, Step Functions, Lambda, and DMS streaming.
Replaced Spark-heavy jobs with a DuckDB + Lambda architecture, cutting compute costs ~40% and processing time ~60% versus the equivalent Glue cluster approach.
Developed REST API endpoints with FastAPI and contributed to the Django backend, owning features end to end from API design to production deployment.
Offloaded analytics from the main application database into a dedicated PostgreSQL data warehouse and integrated ThoughtSpot BI for natural-language querying (NLQ).
Built an AI-enhanced Quarto documentation system covering Step Functions, Lambdas, and custom packages to speed up team onboarding.
Implemented CI/CD with GitHub Actions and infrastructure as code with AWS SAM, CloudFormation, and Terraform.
Python
SQL
PostgreSQL
AWS Glue
AWS Step Functions
AWS Lambda
AWS DMS
DuckDB
FastAPI
Django
ThoughtSpot
GitHub Actions
AWS SAM
CloudFormation
Terraform
Quarto
Business Intelligence Engineer IV · Rackspace Technology
2024 – 2025 · Remote (Mexico City)
Built enterprise ETL automation on BigQuery, created the organization's first Power BI CI/CD pipeline, and led a cost-optimization initiative worth over USD 1,000,000 per year.
Built Python ETL pipelines executing SQL against BigQuery and connecting 10+ enterprise APIs (Azure Monitor SDK, LeanIX OData, ServiceNow, SharePoint, Power BI); the AskHR pipeline ran hourly with zero manual intervention.
Independently designed and shipped an end-to-end Power BI CI/CD pipeline (Python + GitHub Actions + Azure AD service principals) for 90+ dashboards across 70+ workspaces, cutting deployment time from 60 minutes to 2 minutes.
Led BigQuery SQL performance tuning and decommissioning of stale assets, delivering estimated annual savings exceeding USD 1,000,000 and freeing 1,500+ engineering hours per year.
Built LLM-based applications for HR and BI use cases with OpenAI, Gemini, and agentic AI frameworks.
Worked with Databricks, Delta Lake, Unity Catalog, and GCP Dataproc/Dataflow for large-scale processing.
Python
SQL
BigQuery
Dataform
GitHub Actions
Power BI
Azure AD
Databricks
Delta Lake
PySpark
GCP Dataproc
GCP Dataflow
OpenAI API
Gemini API
Tableau
Jira
Senior BI Engineer · Baz Super App
2023 – 2024 · Mexico City
Led cross-cloud financial data streaming and automated marketing analytics on GCP.
Led sensitive financial data replication from AWS to BigQuery through a Kafka streaming pipeline, coordinating security, cloud, and data teams for compliant real-time availability.
Built and automated a Python ETL pipeline ingesting CleverTap marketing events into BigQuery, scheduled with GitHub Actions cron, enabling self-serve campaign reporting.
Implemented data validation and monitoring across cloud environments; worked with Databricks, Dataproc, Dataflow/Apache Beam, Pub/Sub, Cloud Functions, and Looker.
Python
SQL
BigQuery
Kafka
AWS
GCP
Cloud Composer
Pub/Sub
Apache Beam
Databricks
Looker
CleverTap
GitHub Actions
Data Engineer / BI Lead · DiDi Food
2021 – 2023 · Mexico City
Optimized Spark SQL workloads, automated analyst reporting, and led a SQL competency program for the team.
Tuned Spark SQL jobs via Python profiling and choke-point debugging, achieving a 40% reduction in job execution time.
Parameterized and automated SQL queries previously run manually by business users, saving ~100+ analyst hours per quarter.
Maintained data governance across customer-experience analytics pipelines spanning SQL Server and AWS sources (Athena, Glue, S3).
Designed and led a query competency program that raised the team's internal performance score and unlocked additional compute allocation.
Python
SQL
Spark
Presto
Hive
AWS Athena
AWS Glue
AWS S3
Airflow
dbt
Tableau
SQL Server
GitLab
Linux
Data Engineer · DGTIC UNAM
2019 – 2021 · Mexico City
Built the organization's first data warehouse for online lectures and pioneered Spark adoption for MOOC analytics.
Pioneered Spark adoption and built the first data warehouse for online lectures during the pandemic.
Collaborated with IT on database architecture and queried MOOC data from PostgreSQL for Big Data analytics.