Skip to content

Repository files navigation

🍋 citrx

Local-first Apache & Nginx access-log analysis, in your terminal

Stream huge access logs, detect attacks and abuse with deterministic local rules, and explore everything in an interactive TUI.

npm node types license privacy

English · Español


# A single file, a folder, compressed or plain — citrx figures it out
npx @javipm/citrx@latest /var/log/nginx/access.log
npx @javipm/citrx@latest /var/log/nginx/          # a whole folder of logs
npx @javipm/citrx@latest access.log.gz logs.zip   # .gz .br .zip .tar.gz .tgz
cat access.log | npx @javipm/citrx@latest -        # stdin

That command streams the input, validates it, runs ~30 detection rules, and opens a full-screen TUI. No account, no upload, no telemetry.

citrx TUI — summary screen with incident tabs and the global access-log table

🖼️ Screenshots

Summary screen — incident tabs + global access-log table
Summary — incident tabs + indexed access-log table
Incident screen — evidence + related rows
Incident — evidence + related access-log rows
Top values screen
Top values — top IPs, paths, UAs, statuses, params
Filter bar with a query expression
Filter — query language across the log
Terminal report
Terminal report--no-interactive
Self-contained HTML report
HTML report — self-contained, offline

📑 Table of contents


🤔 Why citrx

Access logs quietly hide expensive crawlers, scanner noise, fake bots, SQLi/XSS payloads, POST abuse, and traffic spikes. citrx is built for DevOps, security engineers, and backend developers who need fast answers to:

  • What happened?
  • Which paths, IPs, methods, user-agents, and query params are involved?
  • Which requests should I actually inspect?
  • Which WAF / rate-limit rule would reduce the impact?

The workflow is deliberately offline-first:

1. Run deterministic local analysis   →  no network, bounded memory
2. Explore incidents + raw requests    →  interactive TUI
3. Filter, sort, inspect, select rows  →  small query language
4. Export evidence (TUI: CSV/JSON/TSV; CLI: JSON/Markdown/HTML/terminal)

✨ Features

🌊 Streaming Bounded-memory, line-by-line parsing. Multi-GB logs never load fully into RAM. Top maps keep 20k keys; path stats 8k paths, 64 IPs/path, 256 query variants/path, plus global path-IP/variant budgets (further distinct keys are dropped, not estimated). Dropped-key counters are a lower bound (>=8192) once the fingerprint set is full. RPS histogram keeps 100k occupied seconds plus a bounded overflow set of distinct omitted seconds. Truncation is shown in terminal/Markdown/HTML/TUI.
🧭 Format auto-detect Samples each input, picks apache_common or combined (apache_combined; Nginx combined is the same regex), skips non-access-log files found inside a directory (error_log, xferlog, OS junk) with a warning, and fails only when no input is an access log.
🧩 Custom formats Declarative JSON config with one regex + named fields, validated with zod.
🛡️ ~30 detection rules SQLi/XSS/LFI/SSRF/cmd-injection, recon, fake bots, scanners, DDoS bursts, AI crawlers, POST hotspots, error storms.
🖥️ Full TUI Incident tabs, indexed access-log table, on-demand row loading, top values, request detail, exports.
🔎 Query language AND/OR/NOT, parentheses, field operators, status families, wildcards, per-param filters.
📤 Reports Terminal, JSON, Markdown, and self-contained offline HTML.
📦 Compressed inputs .gz, .br, .zip, .tar.gz, .tgz, folders, and stdin.
🔒 Local-first No telemetry, secrets redacted, temp index deleted on exit.

🚀 Quick start

Run without installing

# npm
npx @javipm/citrx@latest /var/log/nginx/access.log

# pnpm
pnpx @javipm/citrx@latest /var/log/nginx/access.log

# yarn
yarn dlx @javipm/citrx@latest /var/log/nginx/access.log

# bun
bunx @javipm/citrx@latest /var/log/nginx/access.log

Use the @latest tag: npx reuses a cached copy for a bare package name and won't re-check the registry, so without it you may keep running an older version. Alternatively, install globally (below).

Install globally

npm i -g @javipm/citrx
citrx /var/log/nginx/access.log

Common invocations

# Analyze many paths, folders, and compressed files at once
citrx ./logs access.log.gz archive.zip

# Read from stdin
cat access.log | citrx -

# Non-interactive terminal report (CI, pipes, cron)
citrx access.log --no-interactive

# Structured reports
citrx access.log --json
citrx access.log --markdown --out report.md
citrx access.log --html     --out report.html

# Restrict a date range
citrx access.log --since 2026-05-25T00:00:00Z --until 2026-05-25T23:59:59Z

# Force a parser
citrx access.log --format apache_combined

Requirements: Node.js >=22 (developed and tested on 24.15). npx/pnpx handle the rest.

Exit codes make citrx CI-friendly:

Code Meaning
0 Success, no high/critical incidents
1 Execution / configuration error
2 High or critical incidents found

📟 What the output looks like

Running the non-interactive report on a small synthetic log (citrx demo_access.log --no-interactive):

citrx access log analysis

Files: 1
Lines: 72/72
Invalid: 0
Bytes served: 86972
Time range: 2026-05-25T10:00:01.000Z to 2026-05-25T10:05:59.000Z
Peak global RPS: 3 at 2026-05-25T10:03:00.000Z
Formats: apache_combined

Top IPs
      60  8.8.4.4
       4  198.51.100.23
       3  45.83.66.12
       2  192.0.2.55
...

Known AI bots
       3  GPTBot ips=1 paths=1 robots=no

Security incidents (attacks)
  critical 100  SQL injection payload count=1
       ip: 198.51.100.23
       /index.php
       sample: /index.php?id=1+AND+SLEEP(5)
  critical  95  Known scanner user-agent count=4
       ip: 198.51.100.23
  critical  90  Sensitive file probe count=2 2XX_HIT
       ip: 198.51.100.23
       /.env
       /.git/config
  high      85  Known scanner user-agent count=2
       ip: 192.0.2.55

2XX_HIT means the payload or probe received at least one 2xx response — a possible successful reply worth inspecting, not proven compromise.


🧰 CLI reference

Usage: citrx [options] <paths...>

Options:
  --json                  Write machine-readable JSON output.
  --markdown              Write Markdown output.
  --html                  Write a self-contained HTML report.
  --out <path>            Write report output to a file.
  --no-interactive        Print the terminal report instead of opening the TUI.
  --format <format>       auto, apache_common, apache_combined,
                          nginx_combined (same combined regex),
                          or custom:<name>.                   (default: auto)
  --format-config <path>  JSON file with custom access-log formats.
  --top <n>               Limit top lists.                    (default: 20)
  --since <date>          Include entries at or after this date.
  --until <date>          Include entries at or before this date.
  --include <glob>        Include paths matching this glob.
  --exclude <glob>        Exclude paths matching this glob.
  --no-color              Disable colored terminal output.
  --debug                 Print debug details on failure.
  -v, --version           Display the current version.
  -h, --help              Display help for command.

Environment:

  • NO_COLOR=1 — disable color.
  • CITRX_QUIET=1 — silence startup/progress noise for terminal output.

If stdout/stdin are TTYs and no report format is requested, citrx opens the TUI by default. --no-interactive prints the terminal report instead.

--since / --until drop lines whose timestamps cannot be parsed. With no date filter those lines stay in the analysis. TUI timestamp sort is chronological (epoch, including timezone) with row-number tie-break, not log/stream order.


📥 Inputs & formats

Supported inputs

Individual files · folders · stdin (-) · .gz · .br · .zip · .tar.gz · .tgz

ZIP/TAR archives are scanned for candidate log files (access.log, .log, .txt, extensionless logs, .gz, .br). Everything is streamed — full logs are never read into memory. The TUI builds a temporary access-log index under the OS temp dir and removes it on exit.

Built-in formats

  • apache_common — NCSA Common Log Format
  • apache_combined / nginx_combined — NCSA Combined Log Format (identical regex)

Auto-detect reports apache_combined for combined lines. Apache and Nginx combined cannot be distinguished from the line shape alone. --format nginx_combined remains a supported explicit alias of the same parser.

Custom formats

One declarative JSON config, one regex with named groups, validated by zod:

{
  "formats": [
    {
      "name": "pipe",
      "pattern": "^(?<ip>\\S+)\\|(?<timestamp>[^|]+)\\|(?<method>\\S+)\\|(?<target>\\S+)\\|(?<protocol>HTTP/[^|]+)\\|(?<status>\\d{3})\\|(?<bytes>\\S+)\\|(?<userAgent>.*)$",
      "fields": {
        "ip": "ip",
        "timestamp": "timestamp",
        "method": "method",
        "target": "target",
        "protocol": "protocol",
        "status": "status",
        "bytes": "bytes",
        "userAgent": "userAgent"
      }
    }
  ]
}
citrx access.log --format custom:pipe --format-config ./formats.json

Required fields: ip, timestamp, method, target, protocol, status. Optional: bytes, referer, userAgent, host, requestTime, upstreamTime, forwardedFor.

The pattern must use named groups ((?<ip>...)), be anchored with ^ and $, and map every fields.* value to a named group. Numeric capture indices are rejected. Nested-quantifier / ambiguous regexes fail with an actionable error.

host, requestTime, upstreamTime, and forwardedFor are parsed into the internal entry when mapped. They are not yet used by detection rules, TUI columns, or report tables.


🖥️ Interactive TUI

When stdout/stdin are TTYs and no report format is requested, citrx opens a full-screen terminal UI. It's the core product surface, not a debug view.

┌─ citrx ────────────────────────────────────────────────────────────────────┐
│  [ access log ] [ SATURATION ] [ SECURITY ] [ OTHER ]          Tab to switch │
├──────────────────────────────────────────────────────────────────────────────┤
│  #     IP              TIME      MTH  ST   BYTES  PATH                        │
│  3     198.51.100.23   10:01:11  GET  500      0  /index.php?id=1+AND+SLEEP.. │
│  5     198.51.100.23   10:01:12  GET  200   1200  /.env                       │
│  7     192.0.2.55      10:02:00  GET  404      0  /wp-admin/                  │
│ ...                                                                           │
├──────────────────────────────────────────────────────────────────────────────┤
│  f filter   s sort   t top   Enter detail   e export   h help               │
└──────────────────────────────────────────────────────────────────────────────┘
citrx incident screen

Summary screen

Incident area has three tabs (cycle with Tab: access log → SATURATION → SECURITY → OTHER → access log):

Tab Contents
🌊 SATURATION (default) Rate bursts, DDoS, AI crawlers, abusive bots — traffic/resource abuse
🛡️ SECURITY SQLi/XSS/LFI payloads, recon, fake bots, scanner UAs — compromise attempts
🗂️ OTHER Low-signal / noise incidents filtered from the main panels
Tab              switch focus between access log and incident panels
↑/↓              move row            PgUp/PgDn   page through rows
Enter / d        open incident or request detail
f or /           filter access-log rows (Tab cycles example presets)
s or S           open sort menu      t           global top values
Space            select current row  A           select visible rows
e                open export menu (CSV, JSON, TSV only)
r                reset filter, sort, and row selection
h                contextual help overlay (keys + filter syntax)
Esc              cancel the active long-running job, then navigate
q                ask before quit

Sort columns: timestamp (default, descending), ip, status, method, path, bytes. Equal keys always tie-break by row number ascending (stream order). Invalid timestamps sort last in both directions. Selection is capped at 5 000 rows. Incident rows load in buckets of 200.

Incident screen

Evidence + every related access-log line. Rows load on demand by fixed-size buckets, so even huge incidents are responsive immediately. Filtering or sorting a large incident shows background progress in the status bar — press Esc to cancel and revert.

↑/↓ · PgUp/PgDn  navigate            Enter / d   open request detail
t                top values for this incident (computed from full row set)
f · s/S          filter · sort       Space · A   select row · all / visible page
e                export              r           reset filter + selection
Esc              cancel active job, then back
b                back to summary

A selects every matching incident row when the total is ≤ 5 000; above that it only selects the visible page. Manual Space selection is also capped at 5 000. Filter/sort/top/export show progress and Esc cancels them.

Top values · request detail · export

  • Top values (t): top IPs, paths, user-agents, statuses, query params, and param values. Respects the active filter. Enter applies a filter from a value.
  • Request detail (Enter/d): full source, timestamp, IP, method, status, bytes, path, target, user-agent, and raw line with wrapping.
  • Export (e): CSV / JSON / TSV (not Markdown/HTML). Summary with no selection exports the full filtered access-log result; a selection exports only those rows. Incident export streams index rows by chunks to a unique temp file and atomically replaces the destination when finished. Esc aborts a running export.

Long-running filter/sort/top/export operations always show a loading state — the app never looks frozen — and Esc consistently cancels the active operation before navigating.


🔎 Filtering

Filters work on the global access log, incident rows, and top-value drill-downs. Case-insensitive, with a small query language:

  • plain text searches across IP, time, method, path, target, status, bytes, UA, raw line
  • adjacent terms mean AND; explicit AND, OR, |, parentheses, and !/NOT
  • : = contains, = = exact, != = negated match
  • >, >=, <, <= for status, bytes, line
  • status families: status:2xx, status:3xx, status:4xx, status:5xx
  • anchored wildcards: ip:66.249.*
  • quoted values for spaces/symbols: ua:"Googlebot/2.1"
  • URL-encoded values are decoded before matching
method:POST status:200 url:*admin*
(method:POST OR method:PUT) status:2xx
(status:403 | status:404) !ua:*Googlebot*
ip:66.249.* bytes>50000
status:5xx path:/checkout
method!=GET status>=400
param:q                # any request with a q parameter
param:q=*select*       # q value contains "select"
param:*=*sleep*        # any param value contains "sleep"
raw:"union select"
source:access.log line>=10000 line<20000

Fields: ip, method, status, path, target, url, ua, bytes, param, query, source, line, time, raw

Aliases: url→target, timestamp→time, userAgent→ua, st→status, ln→line, src→source, qs→query, mth→method, params→param

Bare text is great for quick hunting — googlebot checkout requires both words somewhere in the searchable line.


📊 Reports

Format Flag Notes
Terminal --no-interactive (or non-TTY) Colored summary + incidents
JSON --json Machine-readable, typed report model
Markdown --markdown Great for tickets / PRs
HTML --html Self-contained, offline, no external resources

Use --out <path> to write to disk. HTML reports are a single offline file: inline CSS and JS, no external network resources, all data HTML-escaped, client-side table filter, click-to-sort tables, executive summary, timeline, incidents, paths, IPs, user agents, plus payloads and suggested actions when those exist. Print/PDF CSS hides the filter toolbar.


🛡️ Detection rules

Every incident carries a kind that drives its TUI panel:

Kind Panel Examples
compromise 🛡️ SECURITY SQLi/XSS/LFI payloads, recon, fake bots, scanner tools
saturation 🌊 SATURATION DDoS bursts, AI crawlers, abusive crawlers, POST hotspots
noise 🗂️ OTHER Low-signal patterns unlikely to need immediate action
Payload & recon rules
ID prefix Category Kind Meaning
sqli: sql_injection compromise union select, sleep/benchmark, encoded SQL
xss: xss compromise script/browser execution indicators
lfi_rfi: path_traversal compromise traversal, LFI/RFI, php://filter, sensitive paths
ssrf: ssrf compromise localhost, metadata IPs/hosts, callback-like params
command_injection: command_injection compromise shell metacharacters + command indicators
recon_sensitive_file: recon compromise probes for .env, .git, backups, dumps
rare_method: http_anomaly noise uncommon methods (CONNECT, TRACE, OPTIONS)
auth_abuse: auth_abuse compromise credential stuffing / brute force on a login endpoint

Payload incidents are grouped by attacker IP (one incident per IP). Scoring by response outcome:

  • any 2xx → SECURITY, critical/100 + 2XX_HIT (payload landed)
  • any 5xx → SECURITY, critical/90
  • only blocked/redirected → OTHER noise (context, not proven impact)
  • requested from a verified Googlebot/Bingbot IP → OTHER noise: the crawler is re-fetching a poisoned indexed URL, so the finding is about the URL, not the IP

auth_abuse: fires on a login endpoint receiving a burst of mostly-failing attempts, either spread across many IPs (credential stuffing) or concentrated on one (brute force). It deliberately sets no 2XX_HIT: most frameworks answer a failed login with 200.

recon_sensitive_file needs ≥2 successful responses or a 10% success ratio to avoid flagging ordinary 404 scanners — except when a high-value target (phpinfo.php, .env, .git/config, wp-config.php.bak, server-status, a .sql dump…) actually returns content, which escalates on its own: one hit out of hundreds of failures is still a leak.

That escalation is withdrawn (servedBodyMatchesGenericPage) when the response weighs exactly what the site serves on ordinary paths — many sites answer any unknown path with 200 and the homepage, and a real phpinfo dump does not weigh what the homepage weighs.

Aggregate path, rate / DDoS, error-storm rules
ID prefix Category Kind Meaning
abusive_crawl: abusive_crawling saturation/noise served path pressure or distributed crawling on a non-entrypoint path
query_explosion: abusive_crawling noise one path with many query variants
post_hotspot: post_hotspot noise endpoint with unusually many POSTs
ddos_rps_burst_single_ip: ddos saturation one IP exceeds per-second RPS for consecutive seconds
ddos_global_rps_spike ddos saturation global RPS over baseline for consecutive seconds
http_head_flood: ddos saturation one IP with high ratio + peak of HEAD requests
ddos_distributed_subnet: ddos saturation IPv4 /24 or IPv6 /48 over RPS + unique-IP thresholds
http_4xx_storm: http_anomaly noise one IP, many 4xx in adjacent minute buckets
http_5xx_storm: http_anomaly saturation one IP, many 5xx in adjacent minute buckets
ddos_sustained_ip_flood: ddos saturation one IP at a high per-minute rate against very few URLs
server_capacity_distress http_anomaly saturation site-wide 503/504/507/508 responses: capacity ran out under load
Bot & scanner rules
ID prefix Category Kind Meaning
ai_scraper_known: ai_scraper saturation/noise known AI crawler/assistant UA, grouped by bot
scanner_ua_known: scanner compromise known scanner/offensive tooling UA
scanner_signature_paths: scanner compromise one IP touches many fingerprint paths
single_ip_path_explosion: abusive_crawling saturation one IP > 10 unique paths/minute sustained
ua_rotation_same_ip: http_anomaly noise one IP, many UAs and peak RPS ≥ 5
fake_bot_googlebot: fake_bot compromise claims Googlebot but IP outside published ranges
fake_bot_bingbot: fake_bot compromise claims bingbot but IP outside published Bing ranges
fake_bot_campaign: fake_bot compromise many IPs sending the same forged Googlebot/bingbot UA (replaces the per-IP rows above 5 sources)
fake_ai_bot: fake_bot compromise AI-crawler UA sent from outside the ranges its operator publishes

abusive_crawl: and auth_abuse: evidence names the heaviest source IPs (topIps, topIpShare) and the heaviest /24-/48 (topSubnet, topSubnetShare), so a path saturated by one client is distinguishable from genuinely distributed traffic.

Notes: single_ip_path_explosion needs pathsPerMinute ≥ 10 (asset-heavy page loads don't trigger it). Server distress is measured as a share of a path's requests (≥2%, with a floor of 100 errors), not as an absolute count: on a million-request URL a hundred 5xx is the background rate every busy endpoint accumulates, so a tally alone mislabelled exactly the highest-volume paths. abusive_crawl enters SATURATION with real served volume + a served-per-minute peak, or with high-peak query churn that still serves some expensive responses even when most attempts are blocked. fake_bot_* needs ≥10 requests. Verified Googlebot/Bingbot IPs are excluded from all bot/scanner detections, and so are loopback/private addresses (127.0.0.0/8, RFC1918, link-local): those identify the server's own proxy hop, not a remote client. AI-crawler identity is checked by IP only for operators that publish their ranges (OpenAI's GPTBot / OAI-SearchBot / ChatGPT-User, Perplexity's PerplexityBot / Perplexity-User). Every ai_scraper_known: incident carries ipVerifiable, so a self-declared UA with no published ranges is never presented as verified.

ai_scraper_known: reaches SATURATION on any one of sustained path fan-out, bot-induced 5xx, or taking a dominant share of total traffic — a crawler hammering a single faceted URL costs the same whether or not it spreads across paths.

Refresh the bundled Googlebot/Bingbot IP-range snapshots with:

pnpm run update-bot-ranges

🎯 Scoring

Each incident has kind, severity, score (0–100), typed evidence, redacted samples, and successful?.

Score Severity
0–24 info
25–49 low
50–74 medium
75–89 high
90–100 critical

Post-processing multipliers:

  • +10 when the same evidence.ip appears in ≥2 incidents (correlated attacker)
  • +15 when a pattern persists ≥30 min (persistence bonus)
  • −10 for moderate AI crawlers that requested robots.txt

Persistence bonus does not apply to ai_scraper_known:* — AI crawlers run for weeks, so duration alone isn't a signal. Panels sort by kind weight (compromise → saturation → noise), then by score descending.


🔒 Security & privacy

  • Local analysis first — no network call during analysis.
  • No telemetry, ever. (If ever added, strict opt-in only.)
  • Secrets redacted in URL/query values: token, _token, sid, session, password, passwd, key, secret, jwt, auth, authorization
  • HTML output escaped; log content is never executed.
  • Temp TUI index files are deleted on exit.

Treat logs, exported JSON, paths, IPs, and route names as sensitive customer data — keep them out of public commits.


🛠️ Development

pnpm install
pnpm run typecheck
pnpm lint
pnpm run format:check
pnpm test
pnpm run build

# run from source against a fixture
pnpm run dev -- examples/your.log
pnpm run dev -- examples/your.log --json

Project layout:

input/    path discovery, stdin, compressed/archive readers
parser/   format detection, parser registry, built-in + custom parsers
analysis/ streaming aggregation, behavior tracking, incident match sets
rules/    deterministic request/path rules and scoring
run/      temporary run workspace and access-log index
tui/      Ink screens, hooks, filters, tables, overlays
report/   terminal, JSON, Markdown, HTML renderers

Stack: TypeScript (ESM) · commander · ink + React · zod · picocolors · Vitest.


📄 License

MIT © javipm

Built for people who read their access logs. 🍋

About

Local-first CLI/TUI for Apache & Nginx access log analysis. Streams large logs, detects SQLi/XSS/LFI, DDoS, fake bots, AI scrapers and more — classified as compromise vs saturation. Interactive terminal UI with filtering. No telemetry.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages