dar4datascience
Menu

Senior Full-Stack Data Engineer · Mexico City

Daniel Amieva Rodriguez

Senior Full-Stack Data Engineer | Data Lakes from Scratch | Multi-Cloud (AWS, GCP, Azure) | Backend-to-Analytics Delivery

I own the whole data path: backend modules that emit the data, third-party API integrations that enrich it, the lake or warehouse that stores it, and the BI and AI layer that serves it. I architect data lakes from scratch or take over and optimize existing ones, across AWS, GCP, and Azure.

Key results

USD 1M+

annual savings

60→2 min

Power BI deploys

~40%

compute cost cut

6+

years experience

Companies

Tech stack

About

The short version

Daniel Amieva Rodriguez is a Senior Full-Stack Data Engineer based in Mexico City with 6+ years of experience delivering data platforms end to end across AWS, GCP, and Azure. He architects data lakes from scratch (a medallion lake on AWS Glue, Step Functions, Lambda, and DMS at TeamStation AI) and takes over existing platforms to optimize them (BigQuery tuning and asset decommissioning worth USD 1M+ per year at Rackspace Technology; 40% faster Spark SQL at DiDi Food). He works the full stack of the data path: FastAPI and Django backend modules that send application data into the lake, integrations with 10+ third-party services and APIs (Azure Monitor, ServiceNow, LeanIX, SharePoint, Power BI, CleverTap, ThoughtSpot), cross-cloud streaming (AWS to BigQuery via Kafka), CI/CD and infrastructure as code, and LLM/MCP integration so business users and AI agents can query governed data. Previous roles: Rackspace Technology, Baz Super App, DiDi Food, and DGTIC UNAM.

What I bring

What I bring

Architects data lakes from scratch

Architects data lakes from scratch (medallion architecture on AWS serverless) and optimizes inherited platforms (BigQuery, Spark, Databricks).

Full-stack ownership

Full-stack ownership: builds the FastAPI/Django backend modules that emit data, the pipelines that move it, and the BI/AI layer that serves it.

Connects third-party services into the platform

Connects third-party services into the platform: 10+ enterprise APIs at Rackspace (Azure Monitor SDK, LeanIX OData, ServiceNow, SharePoint, Power BI), CleverTap at Baz, ThoughtSpot at TeamStation AI.

Multi-cloud delivery

Multi-cloud delivery: AWS (Glue, Lambda, Step Functions, DMS, SAM), GCP (BigQuery, Dataform, Dataproc, Dataflow, Pub/Sub), Azure (AD, Monitor, Fabric), including cross-cloud streaming from AWS to BigQuery with Kafka.

Self-starter with measurable outcomes

Self-starter with measurable outcomes: USD 1M+ annual savings, deployments cut from 60 to 2 minutes, ~40% compute cost and ~60% processing-time reductions.

Current role

Where I work now

Senior Data Engineer · TeamStation AI

May 2025 – Present · Remote (Mexico City) — consultancy for international clients

Architected an enterprise medallion data lake from scratch on AWS serverless and shipped self-service, natural-language analytics for business stakeholders.

See all experience →

Selected projects

Things I've built

Equal Earth Poipoi

R Shiny dashboard that compares the area of any two countries under Mercator (EPSG:3395) and Equal Earth (EPSG:8857) projections, overlays each country's Mercator outline on its true shape to quantify distortion (e.g. Greenland ~16×), and compares Latin America as a block against the United States. Projections are re-centred on each country's meridian to avoid antimeridian splits.

Demo SNIIM Mexico

End-to-end retrieval example for Mexico's SNIIM market-price system (Secretaría de Economía): discovers available markets, builds date-range queries, downloads the HTML results and parses them into a clean pandas DataFrame with optional CSV export. Documented as a Quarto site.

Catálogo Películas Biblioteca Vasconcelos

Data pipeline that turns the Biblioteca Vasconcelos film collection, published only as PDFs, into a searchable catalog: compares regex, Camelot and hybrid PDF table-extraction methods, enriches titles with TMDB/OMDb metadata and director filmographies, routes hard matches to fuzzy review, and publishes an interactive Quarto + ObservableJS site via GitHub Actions.

DuckDB Eurostat MCP Server

Model Context Protocol server that answers natural-language questions over Eurostat data by translating them to SQL and executing them with the DuckDB Eurostat extension (filter pushdown). Supports Anthropic, OpenAI, Azure OpenAI or local Ollama models, plus dataset discovery and schema inspection.

All projects →

Skills

Toolbox

Languages

  • Python
  • SQL
  • PySpark
  • R
  • JavaScript
  • Bash

AWS

  • Glue
  • Lambda
  • Step Functions
  • DMS
  • S3
  • Athena
  • SAM
  • CloudFormation

GCP

  • BigQuery
  • Dataform
  • Dataproc
  • Dataflow / Apache Beam
  • Pub/Sub
  • Cloud Functions
  • Cloud Composer
  • Cloud Storage
  • Cloud Logging

Data Platforms & Processing

  • Databricks
  • Delta Lake
  • Unity Catalog
  • Spark
  • DuckDB
  • Kafka
  • Presto
  • Hive
  • PostgreSQL
  • Microsoft Fabric
  • NoSQL databases

Architecture & Modeling

  • Medallion architecture
  • ETL/ELT pipeline design
  • Streaming pipelines
  • Dimensional modeling (star / snowflake schema)
  • Data governance

DevOps & IaC

  • GitHub Actions
  • GitLab CI
  • Terraform
  • Docker
  • Git / GitHub / GitLab / Bitbucket
  • Jira
  • Confluence

BI & Analytics

  • Power BI
  • Tableau
  • Looker
  • ThoughtSpot (NLQ)
  • Quarto

AI Engineering

  • LLM integration (OpenAI, Gemini)
  • Model Context Protocol (MCP)
  • Agentic AI frameworks
  • Structured outputs
  • Natural-language querying

Backend

  • FastAPI
  • Django
  • REST APIs
  • OOP

FAQ

Frequently asked questions

Who is Daniel Amieva Rodriguez?

Daniel Amieva Rodriguez is a Senior Data Engineer from Mexico City, Mexico, with 6+ years of experience building data platforms on AWS and GCP. He currently works at TeamStation AI and previously held data and BI engineering roles at Rackspace Technology, Baz Super App, DiDi Food, and DGTIC UNAM.

What does Daniel Amieva Rodriguez do?

He is a full-stack data engineer: he architects cloud data lakes and warehouses from scratch or optimizes existing ones, builds the backend modules (FastAPI, Django) that send application data into the lake, integrates third-party services and APIs, runs Python and SQL ETL/ELT pipelines with CI/CD across AWS, GCP, and Azure, and integrates LLM and MCP tooling so business users and AI agents can query governed data.

Is Daniel Amieva Rodriguez a full-stack data engineer?

Yes. At TeamStation AI he built a medallion data lake from scratch on AWS (Glue, Step Functions, Lambda, DMS), developed FastAPI endpoints and Django backend features that feed it, and integrated ThoughtSpot for natural-language BI. At Rackspace Technology he connected 10+ enterprise APIs into BigQuery pipelines and shipped a Power BI CI/CD pipeline. At Baz Super App he led AWS-to-BigQuery streaming with Kafka.

Can Daniel Amieva Rodriguez build a data lake from scratch?

Yes. At TeamStation AI he architected a ground-up medallion data lake (bronze/silver/gold) with custom Python packages on AWS Glue, Step Functions, Lambda, and DMS streaming, replacing Spark-heavy jobs with DuckDB + Lambda for ~40% lower compute cost and ~60% faster processing. At DGTIC UNAM he built the organization's first data warehouse.

Does Daniel Amieva Rodriguez work in multi-cloud environments?

Yes. He has delivered production work on AWS (Glue, Lambda, Step Functions, DMS, Athena, SAM), GCP (BigQuery, Dataform, Dataproc, Dataflow, Pub/Sub, Cloud Composer), and Azure (Azure AD, Azure Monitor, Microsoft Fabric), including cross-cloud replication of financial data from AWS to BigQuery via Kafka at Baz Super App.

Is Daniel Amieva Rodriguez a data engineer?

Yes. Daniel Amieva Rodriguez is a Senior Data Engineer. His official titles have included Senior Data Engineer (TeamStation AI), Business Intelligence Engineer IV (Rackspace Technology), Senior BI Engineer (Baz Super App), and Data Engineer / BI Lead (DiDi Food).

What technologies does Daniel Amieva Rodriguez specialize in?

Python, SQL, PySpark, AWS (Glue, Lambda, Step Functions, DMS, SAM), GCP (BigQuery, Dataform, Dataproc, Dataflow), Databricks and Delta Lake, DuckDB, Kafka, PostgreSQL, GitHub Actions, Terraform, and LLM/MCP integration.

Where is Daniel Amieva Rodriguez based?

Mexico City, Mexico. He works remotely with international clients and is fluent in English and a native Spanish speaker.

What is Daniel Amieva Rodriguez's education?

A Bachelor's degree in Economics (2015–2021) and a Specialization in Data Science (2019–2020), both from the Universidad Nacional Autónoma de México (UNAM).

What measurable results has Daniel Amieva Rodriguez delivered?

Estimated annual savings exceeding USD 1,000,000 from BigQuery optimization and asset decommissioning at Rackspace Technology; Power BI deployment time cut from 60 to 2 minutes; ~40% compute cost and ~60% processing-time reduction on AWS at TeamStation AI; 40% faster Spark SQL jobs at DiDi Food.

How can I contact Daniel Amieva Rodriguez?

Email danielamieva@dar4datascience.com, or connect on LinkedIn at linkedin.com/in/dar-4-ds. Code is on GitHub at github.com/dar4datascience.

What is dar4datascience?

dar4datascience is the online handle of Daniel Amieva Rodriguez (D.A.R. for data science). It is used for his GitHub account, this website, and his data engineering projects.

Need a data lake built or fixed?

I design, build, and optimize cloud data platforms end to end — from backend emitters to BI and AI layers.