Interactive Learning Course | Home Works & Quiz | Fall 2021 | Prof. Majid Nili
-
Updated
Feb 24, 2022 - Jupyter Notebook
Interactive Learning Course | Home Works & Quiz | Fall 2021 | Prof. Majid Nili
Robust contextual dueling bandits with post-serving context, delayed feedback, and adversarial corruption (RLHF / preference learning) — ICML 2026
c++ implementation of algorithms for solving regret-minimising database queries
C++ implementation of Multi-Armed Bandits (Gaussian and Bernoulli)
A python implementaion of Counterfactual Regret Minimization using numba
This repository contains code for the paper "Non-monotonic Resource Utilization in the Bandits with Knapsacks Problem".
POMRL: No-Regret Learning-to-Plan with Increasing Horizons [TMLR 2023]
Create a platform that recommends sustainable farming practices to farmers based on their specific location, soil type, crop choice, and climate conditions. Incorporating data on sustainable agriculture methods could help in increasing crop yield, reducing environmental impact, and promoting biodiversity.
Paper implementation of Sequential Learning for Multi-Channel Wireless Network Monitoring With Channel Switching Costs
A visualization of a Regret Minimization Learning algorithm for Two Person games, but Avatar themed! 15-251 Fall 2020 Project
Source code for Regret synthesis for two-player turn-based game played on graphs - ICRA 22
Block-level adaptive compression using LinUCB contextual bandit routing. Outperforms LZ4 by 20.4%, LZMA by 9.9%.
A regret minimization approach to training Generative Adversarial Networks (GANs). This was my project in the "Algorithms and Optimization for Big Data" course.
Project on preference learning - ENSAE ParisTech
This repository contains several implementations of multi-armed bandit (MAB) agents applied to a simulated cricket match where an agent selects among different strategies with the goal of maximizing runs while minimizing the risk of getting out.
CFR and equilibrium-computation research: notes, benchmarks and methodology, applied to 5-card PLO
Probabilistic Future Video Frame Prediction using Generative Adversarial Networks by employing a regret minimization strategy for training GANs.
Add a description, image, and links to the regret-minimization topic page so that developers can more easily learn about it.
To associate your repository with the regret-minimization topic, visit your repo's landing page and select "manage topics."