We measure that engines disagree. We do not yet say why each disagreement happens, and that is the difference between an observation and a finding.
ADR-0010 requires every divergence to be labelled one of two things:
- Specification difference — the engines made different, defensible design decisions. radon's Halstead operator set is this: narrow, undocumented, but not wrong.
- Bug — behaviour the engine's own maintainer would not defend. lizard's
-ENS counter leak is this.
The label changes what should be done about it. A specification difference gets documented; a bug gets an upstream issue.
The task
For each metric in frame.divergence_summary():
- Find its extreme cases in the corpus.
- Read both engines' implementations for that metric.
- Write down which of the two categories it falls into, and the mechanism.
Why it is the highest-value thing here
This is the part a competitor cannot copy by reading our code, because the value is in the empirical work and in having read the implementations. It is also what turns the project from a useful wrapper into something publishable.
Currently unclassified: halstead_volume, halstead_difficulty, halstead_effort, halstead_length, halstead_vocabulary, cognitive_complexity_mean, cyclomatic_complexity_mean, function_count, class_count, lloc, sloc, dead_code.
Pick one and work it through — a single well-argued classification is a real contribution.
We measure that engines disagree. We do not yet say why each disagreement happens, and that is the difference between an observation and a finding.
ADR-0010 requires every divergence to be labelled one of two things:
-ENScounter leak is this.The label changes what should be done about it. A specification difference gets documented; a bug gets an upstream issue.
The task
For each metric in
frame.divergence_summary():Why it is the highest-value thing here
This is the part a competitor cannot copy by reading our code, because the value is in the empirical work and in having read the implementations. It is also what turns the project from a useful wrapper into something publishable.
Currently unclassified:
halstead_volume,halstead_difficulty,halstead_effort,halstead_length,halstead_vocabulary,cognitive_complexity_mean,cyclomatic_complexity_mean,function_count,class_count,lloc,sloc,dead_code.Pick one and work it through — a single well-argued classification is a real contribution.