Pinned Loading
-
judgecliff
judgecliff PublicWhich judges survive optimization pressure? Exploitability of image-generation QC judges (rules, API VLMs, open VLMs, trained reward models) under BoN, DPO, and SRPO attacks.
Python
-
llama-chess-uci
llama-chess-uci PublicFull fine-tune of Llama 3.1 8B to play legal UCI chess. 45 to 98 percent legal moves, median centipawn loss 255 to 55, Stockfish-verified.
Python
-
research-agent
research-agent PublicAutonomous deep-research agent with a model-driven control loop and a cross-provider eval panel
Python
-
soho-coop-rag
soho-coop-rag PublicRAG Q&A over every co-op building in NYC's SoHo historic district, from public data, with a hybrid SQL/vector router and a control-group evaluation.
Python
If the problem persists, check the GitHub status page or contact support.