Skip to content
View thylinao1's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report thylinao1

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. bayes-cot-faithfulness bayes-cot-faithfulness Public

    Bayesian causal-mediation framework for measuring whether a language model's chain-of-thought causes its answer, treating the reasoning as a mediator and reporting a posterior over the direct and i…

    Python 1

  2. Olist-Bayesian-AB-Testing Olist-Bayesian-AB-Testing Public

    Hierarchical Bayesian A/B testing on the Olist Brazilian e-commerce dataset. DuckDB SQL pipeline feeding PyMC; a difference-in-differences design flips the sign of the naive on-time-delivery result.

    Python 1

  3. car-bench-track1-submission car-bench-track1-submission Public

    Track 1 (Open) entry to the CAR-bench Challenge at IJCAI-ECAI 2026. Placed 2nd of 21 teams on the hidden test set with 66.7% Pass^3 and 86.7% Pass@3. Adds a live tool-schema guard, temperature-zero…

    Python

  4. egress-receipts egress-receipts Public

    Egress receipts for AI evaluation sandboxes: one signed file per run that an outsider verifies offline. Apart x CeSIA AI Incident Response Sprint 2026, Track 1.

    TypeScript

  5. gauntlet gauntlet Public

    Autonomous red-team for AI applications. Fires attack probes at a chat app, scores it against the OWASP LLM Top 10 (2025), then re-runs the same probes with a runtime guard on. Bundled demo is offl…

    TypeScript

  6. lm-kbc-2026 lm-kbc-2026 Public

    Entry to the LM-KBC 2026 challenge on knowledge base construction from language models. Scored 0.7060 all-rows macro-F1 on the official test set, leading on five of the six relations. Final standin…

    Python