Software engineer building trustworthy, reproducible, and inspectable AI systems.
I focus on the engineering boundary between AI capability and operational trust: evidence, provenance, deterministic evaluation, privacy-aware debugging, and failure-resistant developer tools.
| Project | What it demonstrates |
|---|---|
| CorpusSeal | Evidence-first benchmark contamination and dataset integrity auditing with deterministic exact/near matching, SARIF, HTML, and GitHub Actions |
| BidiFence | Deterministic RTL/i18n conformance checks for Playwright with SARIF, baselines, and Arabic fixtures |
| TraceSift | Offline causal diagnosis and privacy-safe regression fixtures for AI-agent traces |
| VeriTrace | Deterministic conformance and replay testing for agent governance |
| ML ProofLedger | Portable, verifiable evidence manifests for machine-learning runs |
| Mizan | Evidence-backed Arabic claim verification with abstention and reproducible evaluation |
AI reliability and observability, benchmark and dataset integrity, OpenTelemetry-compatible trace contracts, reproducible ML evaluation, privacy-preserving artifacts, policy-as-code, Python tooling, API design, test architecture, and open-source maintenance.
Python · TypeScript · FastAPI · pytest · GitHub Actions · OpenTelemetry · Docker · PostgreSQL · React
I prefer small stable contracts over opaque integrations, fail-closed behavior over optimistic guesses, local-first workflows where sensitive data is involved, and documentation that states limitations as clearly as capabilities.
The best way to collaborate is through GitHub Issues and Discussions on the relevant project.