Machine Learning Engineer · Lead Data Scientist · AI Strategist
Dr.-Ing. in Computer Science with 15+ years of experience across applied machine learning, AI, experimentation, and large-scale data systems.
I build production-oriented ML and analytics systems from research and prototyping through architecture, implementation, validation, deployment, and governance. My current technical focus is agentic analytics, distributed data engineering, experimentation, and reproducible ML systems.
| Project | Focus |
|---|---|
| agentic-data-analyst | Guarded agentic analytics: natural language → dbt metadata → typed IR → validated PySpark execution |
| spark-data-quality | Typed, Spark-native data-quality validation and profiling with shared aggregation planning |
| experimentation-toolkit | A/B testing, statistical inference, SRM diagnostics, power analysis, and multiple-testing correction |
| spark-asammdf | Apache Spark DataSource V2 connector for distributed analysis of ASAM MDF4 measurement data |
| retro-speedlab-core | Recurrent PPO + RND reinforcement-learning engine with vectorized environments, resumable checkpoints, and telemetry |
| retro-speedlab | Cookiecutter scaffold for reproducible Stable Retro reinforcement-learning projects with tested project generation and CI |
ML & AI: PyTorch · scikit-learn · Spark ML · LLMs · AI agents · MCP · reinforcement learning
Data & distributed systems: PySpark · Apache Spark · Databricks · Airflow · SQL · data quality · metadata · lineage
Software engineering: Python · Scala · FastAPI · REST · Docker · GitHub Actions · CI/CD · testing · reproducibility
Cloud: Azure · AWS · GCP
I started working on applied machine learning and behavioral analytics at TU Dresden, where I completed my doctorate in Computer Science.
Since then, I have worked across research, healthcare, digital products, and automotive. My recent work has centered on production ML and large-scale automotive analytics: distributed PySpark pipelines, semantic data abstractions and metadata indexing, ML/AI-assisted analytics, reproducible delivery, monitoring, and governance.
The repositories above use public or synthetic data and are independent portfolio projects; they do not contain proprietary employer or client code or data.

