Skip to content
View sjarmak's full-sized avatar

Block or report sjarmak

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sjarmak/README.md

benchmarks + evals: codeprobe · CodeScaleBench · EnterpriseBench · migration-evals · agent-diagnostics · mg-ax
gas city: gascity · gascity-packs · gascity-dashboard   research: agent-code-authorship · mem · agent-oriented-architecture · GEO_public
agent tooling: livedocs · tom-swe · coding-agent-workflows · brains · code-intelligence-digest
scix + search: scix-agent · nls-finetune-scix   play: website · WheelOfFortune · embertide

readout refreshed from real commit activity by generate_readout.py  ·  sjarmak.ai

Pinned Loading

  1. gascity gascity Public

    Forked from gastownhall/gascity

    Orchestration-builder SDK for multi-agent coding workflows

    Go

  2. gascity-dashboard gascity-dashboard Public

    Forked from gastownhall/gascity-dashboard

    Development fork of gastownhall/gascity-dashboard: an editorial-typographic ambient dashboard for a single Gas City operator. Changes are staged here and land upstream.

    TypeScript

  3. EnterpriseBench EnterpriseBench Public

    Benchmark for coding agents on large multi-repo enterprise codebases: 112 tasks across 10 task types covering cross-repo dependency tracing, incident investigation, and non-patch artifacts. WIP; re…

    Python

  4. scix-agent scix-agent Public

    Agent-navigable knowledge layer over the NASA ADS/SciX corpus: 32.4M papers and 299M citation edges exposed through a 15-tool MCP server with hybrid search, citation-graph analytics, and LLM entity…

    Python 6 1

  5. codeprobe codeprobe Public

    Turn a repository's merged pull requests into coding-agent evaluations, then measure the whole agent setup (model, tools, retrieval, cost, harness) against them. Python CLI, on PyPI.

    Python 7

  6. mem mem Public

    Agentic memory built from multi-agent orchestration traces (6,691 work items, 874 session transcripts) and benchmarked on the same record: does retained memory improve agent success rate, iteration…

    Python