Skip to content
View datenwissenschaften's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report datenwissenschaften

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Dr. Martin Franke

Machine Learning Engineer · Lead Data Scientist · AI Strategist

Dr.-Ing. in Computer Science with 15+ years of experience across applied machine learning, AI, experimentation, and large-scale data systems.

I build production-oriented ML and analytics systems from research and prototyping through architecture, implementation, validation, deployment, and governance. My current technical focus is agentic analytics, distributed data engineering, experimentation, and reproducible ML systems.

Selected projects

Project Focus
agentic-data-analyst Guarded agentic analytics: natural language → dbt metadata → typed IR → validated PySpark execution
spark-data-quality Typed, Spark-native data-quality validation and profiling with shared aggregation planning
experimentation-toolkit A/B testing, statistical inference, SRM diagnostics, power analysis, and multiple-testing correction
spark-asammdf Apache Spark DataSource V2 connector for distributed analysis of ASAM MDF4 measurement data
retro-speedlab-core Recurrent PPO + RND reinforcement-learning engine with vectorized environments, resumable checkpoints, and telemetry
retro-speedlab Cookiecutter scaffold for reproducible Stable Retro reinforcement-learning projects with tested project generation and CI

Engineering focus

ML & AI: PyTorch · scikit-learn · Spark ML · LLMs · AI agents · MCP · reinforcement learning

Data & distributed systems: PySpark · Apache Spark · Databricks · Airflow · SQL · data quality · metadata · lineage

Software engineering: Python · Scala · FastAPI · REST · Docker · GitHub Actions · CI/CD · testing · reproducibility

Cloud: Azure · AWS · GCP

Background

I started working on applied machine learning and behavioral analytics at TU Dresden, where I completed my doctorate in Computer Science.

Since then, I have worked across research, healthcare, digital products, and automotive. My recent work has centered on production ML and large-scale automotive analytics: distributed PySpark pipelines, semantic data abstractions and metadata indexing, ML/AI-assisted analytics, reproducible delivery, monitoring, and governance.

The repositories above use public or synthetic data and are independent portfolio projects; they do not contain proprietary employer or client code or data.


Website

Pinned Loading

  1. agentic-data-analyst agentic-data-analyst Public

    Guarded agentic analytics: natural language → dbt metadata → typed IR → validated PySpark execution.

    Python 1

  2. experimentation-toolkit experimentation-toolkit Public

    Statistical toolkit for A/B experiments: inference, SRM diagnostics, power analysis, and multiple testing.

    Python 1

  3. retro-speedlab retro-speedlab Public

    Cookiecutter scaffold for reproducible Stable Retro reinforcement-learning projects with tested project generation and CI.

    Python 1

  4. retro-speedlab-core retro-speedlab-core Public

    Recurrent PPO + RND reinforcement-learning engine with vectorized environments, resumable checkpoints, and telemetry.

    Python 1

  5. spark-asammdf spark-asammdf Public

    Apache Spark DataSource V2 connector for distributed analysis of ASAM MDF4 automotive measurement data.

    Jupyter Notebook 1

  6. spark-data-quality spark-data-quality Public

    Typed, Spark-native data-quality validation and profiling with shared aggregation planning.

    Jupyter Notebook 1