Changelog¶
This page shows the repository's
CHANGELOG.md,
included directly so the two can't drift apart.
Changelog¶
All notable changes to PyCodeCommenter will be documented in this file. Format follows Keep a Changelog.
[Unreleased]¶
[2.6.1] - 2026-09-29¶
Added¶
- A Data and Privacy page in the docs (
docs/data-and-privacy.md): what--ai-draftsends, what the secret check leaves out, how consent works, where the hosted service and each provider send code, and where to report a problem. Linked from the README and the FAQ.
Upgrade notes¶
- Everyone is asked for hosted-AI consent again once. The consent notice
now says the hosted service forwards code to Google's Gemini API on a free
tier, where Google may use it to improve its products, so
CONSENT_NOTICE_VERSIONis 3. CI jobs that use--ai-draftwithout--yes-send-code-to-aiwill fail until someone agrees or the flag is added.
Fixed¶
--ai-draftwithout a terminal no longer crashes at the consent prompt. With no consent on file and nothing to read an answer from (CI, closed standard input) it raisedEOFError; it now stops cleanly with "AI drafting needs your consent" and says to run in a terminal or pass--yes-send-code-to-ai.- The validator no longer asks for a
Returns:section on a-> Nonefunction, and allows one on an abstract method whose body only raisesNotImplementedError(or is...). - Placeholder detection matches whole marker words.
TODO,FIXME,XXX,HACKandTBDare matched as whole uppercase words, andDescription ofonly at the start of a line, so ordinary prose (such as "a description of the file") is no longer reported as a placeholder. - A wrapped
Args:description is no longer split into a phantom entry. A deeper-indented continuation line shaped likeword: textwas read as a new parameter; the validator then reported it as documented but missing from the signature, and regenerating a docstring dropped the author's continuation text. It now continues the entry above.
Changed¶
- New default models for Gemini and Anthropic. Anthropic now defaults to
claude-sonnet-5-5(asked for low effort) instead ofclaude-haiku-4-5-20251001, which Anthropic lists for retirement no sooner than 2026-10-15. Gemini now defaults togemini-3.8-flash; if Google says it is unavailable to your key (not found, no access, or still failing after the SDK's retries) the request moves on togemini-3.5-flash-liteand thengemini-2.5-flash, and stays on the one that worked for the rest of the run. A Gemini model you choose with--ai-modelis never replaced. Nothing changes for the DeepSeek or OpenAI-compatible defaults. - OpenAI now defaults to
gpt-6-luna(about $0.1 / $0.5 per million input / output tokens) instead of the flagshipgpt-6-astra($10 / $50), and GPT-6 models are asked for low reasoning effort, since their default is medium and reasoning tokens are billed as output. A one-sentence docstring does not need the flagship. - The hosted-service consent notice is accurate about Google. It now says
the service does not store your code, forwards it to Google's Gemini API,
and that Google handles it under its free-tier terms and may use it to
improve its products. The old wording ("No source code is stored beyond the
time it takes to process each request") implied no copy existed anywhere.
CONSENT_NOTICE_VERSIONis now 3, so everyone is asked again once. - The package's own docstrings are complete: 0 validator findings and 100%
docstring coverage (also counted with
--strict). CI now checks formatting, lint and the package's own docstring quality on every push. - The package ships a
py.typedmarker.
[2.6.0] - 2026-09-27¶
The tool now states everything the code proves, keeps everything the author
wrote, and makes what's left easy to finish: optional AI drafting of the gaps
(hosted, or with your own key), a review command, and a summary at the end
of every run. Requires Python 3.10+.
Also includes follow-up work from a dogfooding audit run against the tool's own codebase
— every numbered finding in that audit is now closed; see
docs/dev-notes/audit-remediation-log.md for the full history.
Changed¶
- README and docs show real output. The README opens with a real
before/after (with and without
--ai-draft) and the new commands. Two examples in the README and docs showed output the tool doesn't produce (invented parameter descriptions; an old TODO description paragraph); they now show the tool's actual output. Statements that the tool never uses AI now say "by default" and explain the opt-in. The AI section now also says how to set an API key (bash and PowerShell;.envfiles are not read), what a run costs, the hosted limit as a number, and that an AI label stays until a person accepts the line inreview; the validation reference gained theai_draftcheck. - NumPy- and Sphinx-style docstrings keep their style. Regenerating
used to convert them to Google style. Now gaps are filled in the
docstring's own convention -- NumPy's dash-underlined
Parameters/Returns/Yields/Raises/Attributessections, or Sphinx's:param:/:type:/:returns:/:rtype:/:raises:/:ivar:fields -- and new docstrings are still Google style. The parser now also reads NumPyYields/Attributesand Sphinx:yields:/:ivar:/:vartype:, which were previously dropped.
Added¶
pycodecommenter reviewsteps through what needs a person aftergenerate: each AI-drafted line (accept, edit, skip), eachTODOgap (fill, skip), and each comment a new docstring now repeats (remove only on an explicit yes; the default keeps it). Only docstring lines and approved comments change, and a file is saved only if it still parses with exactly the same code.--list, or running without a terminal, lists the items and changes nothing.generateends with a summary of what it did: docstrings written, updated (author text kept) or already complete; details taken from the code; AI-drafted lines; docstrings taken from comments; and gaps left -- followed by a suggested next step. Totals cover the whole run for a directory. Printed to stderr, so it never mixes with code on stdout.- A
#comment block above an undocumented function or class becomes its docstring: the first sentence as the summary, the rest as the description. It counts as the author's own text, so AI drafting never replaces it. The comment is left in place -- the tool never deletes code. Notes (TODO), tool directives (noqa,type:), commented-out code and section banners (blocks with a# -----line) are not used, and neither is a comment separated from the definition by a blank line. --ai-draftfills every gap, not just the description. A name-derived summary, parameters with a TODO or type-only description, an undescribed return value, and exceptions without a readable condition are drafted in one request per function, using the hosted service's new/v2endpoint. Author text and facts read off the code are never replaced. Every drafted line carries the(AI-drafted, unreviewed)marker and passes a safety check before it's written (no triple quotes, backslashes, placeholder or marker text). Drafts are kept on later runs, so regenerating costs no further requests.- Bring your own key:
--ai-provider gemini|openai|anthropic|deepseek| openai-compatiblecalls that provider directly with your key (read fromGEMINI_API_KEY,OPENAI_API_KEY, ... or asked for, hidden), through its official SDK installed as an optional extra (pip install "pycodecommenter[gemini]",[openai],[anthropic], or[ai]for all; Python 3.10+). Each provider has a default model (Anthropic:claude-haiku-4-5-20251001; Gemini:gemini-2.5-flash), printed at the start of every run;--ai-modelchooses any other, and--ai-base-urlpointsopenai-compatibleat Mistral, Groq, Ollama, etc. - Daily limit, then your own key: the hosted service now allows 25 drafts per caller per day. The CLI reports how many are left; when they run out mid-run, an interactive run asks whether to continue with your own key from that same function, and a CI run stops cleanly and says how to continue.
- Consent per destination: agreeing to send code to the hosted service
no longer covers a provider you call directly, or vice versa.
--yes-send-code-to-aireplaces--yes-send-code-to-hosted-ai, which still works as an alias. generate --output-dir PATH(§6): writes a fully-documented copy of a directory target's tree toPATH, mirroring each file's relative path, leaving the originals untouched — closing the exact gap that madedocument_folder.pynecessary as a hand-rolled external script for the audit. Mutually exclusive with--inplace; rejected on a single-file target the same way--outputis already rejected on a directory one.- Opt-in module-level docstring generation (§2), for a module that has
none at all:
PyCodeCommenter(include_module_docstrings=True), orgenerate --include-module-docstringson the CLI. Off by default — unlike a missing function/class docstring, a missing module docstring would touch the output of nearly every input (any file/snippet with no module docstring, not just amain.py-shaped real project file lacking documentation entirely), so this stays behind an explicit flag, the same way--inplace/--backupalready do for other consequential behavior. When enabled, the summary is derived from the file's name (viafrom_file) or a neutral placeholder (viafrom_string, which has no filename to go on), plus a realClasses:/Functions:listing of what the module defines when it defines anything — mirroring theAttributes:/Methods:pattern already used for classes. A module with an existing docstring is never touched: unlike function/class docstrings, this never attempts to merge into one, since module docstrings are far more free-form prose and reconstructing one risks corrupting it for no benefit. Raises:entries state the condition when the code states it exactly (newcode_facts.py): araisedirectly under a top-levelifbecomesValueError: If `discount < 0`., one in theelsebranchIf `x in ALLOWED` is false., one in a top-levelexceptIf `OSError` occurs., and an unconditional one in a function with noreturnAlways.. Nested,elif, in-loop, and over-long conditions keep the guess marker, since one clause can't state them exactly.boolandstrreturn types read off return expressions: comparisons,not,isinstance/hasattr-style built-ins andand/orover booleans givebool; f-strings andstrmethods on a string literal (" ".join(...)) givestr. A function with a single boolean return expression getsTrue if `expr`, otherwise False.instead of a guess marker.--ai-draftnow drafts classes too. Before, only functions were sent to the AI, so every classAttributes:entry and every "<Name>class." summary kept its TODO or type-only filler even with a perfect model. A class request carries an outline of the class (header, class-level fields,__init__in full, other methods as signatures); the same rules apply as for functions (only gaps, author text and facts never replaced, every line labelled, a declined value leaves its gap). Providers gaindraft_class_docstring, which declines by default so existing providers keep working. The hosted service needs its new/v2/draft-class-docstringendpoint; against an older deployment classes simply keep their TODOs.- The run summary says why gaps are left. After AI drafting it counts the parts the AI declined to write, the requests that failed, the functions and classes not tried because drafting stopped, and those not sent because their source looks like it holds a secret. It also reports documented arguments removed because their parameter no longer exists.
- Know the cost first, and cap it. For a directory,
--ai-draftcounts the requests it would make (nothing is sent to find out), prints32 files, 121 functions and classes have gaps to draft, and in a terminal asksContinue? [y/N](default no).--max-drafts Ncaps the requests for the whole run. - Source that looks like a secret is never sent. A function or class containing a private-key header, a well-known API key shape, a JWT, a URL with a password, or a literal assigned to a password/token/key-named variable is kept out of AI requests and keeps its deterministic docstring. The consent notices now say what is sent (function source and class outlines, comments included) and that likely secrets are left out; because that wording changed, consent is asked for once more.
- Opt-in strictness.
coverage --strictcounts only docstrings with noTODO(pycodecommenter)placeholder and no unreviewed AI line, so a project of generated stubs no longer reads as 100%.validate --fail-on-todoand--fail-on-ai-draftexit 1 on those issues. Defaults are unchanged. - A rate limit no longer ends a bring-your-own-key run at once. A 429
is waited out once (the response's
Retry-After, 1-60 seconds, or 20 without one, shown on the status line) before drafting stops. reviewcan accept a whole function's AI lines at once. For a function with two or more AI-drafted lines,reviewshows its whole docstring and offersAto accept all of that function's remaining AI lines. It stops at the function boundary, and there is deliberately no file-wide or global accept, so removing the label still means a person read the text.- Live progress while drafting. On a terminal,
--ai-draftshows one status line per request ([3/12] product_service.py - drafting create_product (hosted)) and says when it is waiting out a rate limit, so a slow first hosted request no longer looks like a hang. Nothing is written when stderr is redirected or in CI. - A missing provider SDK is caught first. The SDK is checked before
the consent question and the key prompt, and the error prints the exact
pip installcommand for the Python that is running (the tool never runs pip itself). When the hosted limit hands over to your own key and that provider can't be used, you are asked again until one works or you typeskip.
Changed¶
Methods:is no longer generated for classes. It isn't a standard Google-style section (the validator already reported it as non-standard), every public method carries its own docstring, and the generated entries were guess markers only. An author's ownMethods:section is kept; entries left over from earlier runs are removed.
Fixed¶
*argsand**kwargsdrafts were dropped when a model answered underlayersinstead of*layers; the unstarred name is now accepted as a fallback, so those parameters stop keeping filler text.True ifbool(x), otherwise False.now readsTrue ifx, otherwise False.; the redundant wrapper is not repeated.- A class with no attributes got a docstring ending in a blank line. It is now one line (or summary, description and closing quotes).
- An attribute assigned from a typed parameter was documented as
Any. It takes the parameter's type. - AI setup errors printed to stdout (declined consent, missing key or SDK); they go to stderr like every other AI message.
- Class
Attributes:listed names that are not attributes. An__init__argument was documented even when__init__never stored it (for example one only passed tosuper().__init__), and one stored under another name appeared under both. An argument now counts only if__init__assignsself.<name>. (default: unknown)for defaults the code shows. A default such asPriority.MEDIUMorDATA_FILEis now shown as written; one over 40 characters is left out of the line instead of guessed.Path to the .for a parameter namedpath. It now reads "Path to the file or directory".- A bare
null.paragraph from AI drafts. A drafted value that is onlynull,none,n/a,nilorundefinedis declined, so the gap keeps its TODO instead of receiving junk text. generate <dir> --output-dir <dir>/outprocessed its own earlier output on the next run (out/out/...). The output directory is now excluded from what is collected.- Consent prompts could be invisible, leaving
generatewaiting. With the patched code going to stdout (generate app.py > out.py), the consent question went intoout.pytoo. Prompts and status messages now go to stderr; stdout carries only generated code. - Escape sequences in existing docstrings were rewritten, and could
break the file. Existing docstrings were read as their value, so a
written
\ncame back as a real line break on every run -- and escaped quotes (\"\"\") came back as a real""", ending the docstring early and producing a file that no longer parses. Docstrings are now read from their source text, escapes kept as written, and a raw string keeps itsrprefix. - An author's
__init__summary was replaced with "Initialize the class."; that text is now only used for an__init__with no docstring. - Text from an earlier run was treated as author text, so its TODO
markers and type-only filler (
float value.) could never be improved -- including by--ai-drafton a file generated before it existed. The generator now recognises its own earlier output and regenerates it. - Defaults were stated twice (
Default is 3. (default: 3)); the(default: ...)suffix is now the only mention. raise NotImplementedError(no parentheses) got noRaises:entry. A bare raise of a built-in exception, or of a name following the exception-class convention (ConfigError), is now documented;raise err(an instance) still isn't guessed.- Regeneration overwrote author-written
Raises:,Attributes:andMethods:text. An author's exception and attribute descriptions were replaced with guess markers (or inferred filler) on every run, and a hand-writtenMethods:section was deleted, so running the tool on a fully documented file made it less documented.DocstringParsernow parses Google, Sphinx (:raises X:) and NumPyRaisesentries, and GoogleAttributes:/Methods:, and the generator carries them forward. Exceptions or attributes the author documented but the code doesn't raise/assign directly (propagated exceptions, class constants) are kept. - Name inference matched "count"/"num" inside other words:
discountbecame "Number of dis." andcountry"Number of try.". Matching is now by whole word. - Untyped class attributes were shown as
(any)while the same parameter inArgs:showed(Any). - CRLF files were silently normalized to LF on every generation run
(§3), even on files with zero docstring changes — turning a
documentation PR on a CRLF file (common on projects with Windows
contributors) into a full-file line-ending diff.
from_file/from_stringnow detect and remember the source's newline convention, normalizing to\nonly for internal processing;get_patched_code()restores the original convention at the final output boundary. The CLI's--inplace/--output/--output-dirreads and writes now usenewline=""too, so the before/after comparison isn't fooled by universal-newlines translation and a CRLF file isn't double-translated to\r\r\non write on Windows. - An unexpected exception mid-generation silently discarded a real,
existing docstring (§5), replacing it with the unhelpful literal
"""Error generating docstring."""— adocument_folder.py-style driver script's own[OK]/[FAIL]reporting would never see this, since the failure was per-function, not per-file._generate_function_ docstring/_generate_class_docstringnow fall back to the original docstring when one existed, matching the fallback discipline this file already uses for a libcst parse failure elsewhere. Unobserved in practice before this fix (the audit flagged it as a latent risk, not an encountered bug), so verified by forcing a real exception directly rather than against a naturally-occurring repro. __init__was the one function still getting fixed boilerplate description text on every constructor, regardless of what the class does ("Initialize a new instance.", ten times across this project's own source alone) — inconsistent with the "no placeholder paragraph" principle already applied to every other function earlier in this thread.__init__'s description now follows the same rule as everywhere else; only the summary line ("Initialize the class.") stays fixed, since the generic name-derived fallback would produce"Init."for__init__, which reads worse than what it would replace.- A
@property's getter/setter/deleter were listed three times under one name in a class'sMethods:section (§1.8), with no indication which was which. TheMethods:list was built from every non-private method in the class body with no deduplication — but a setter/deleter's decorator (<property_name>.setter/.deleter) is only valid Python when it rebinds the exact same name as the property, so any duplicate name here is always one property's accessor trio, never two distinct methods. Fixed: the method-name list now dedupes viadict.fromkeys(), preserving first-seen order. - The validator false-flagged its own honest output on
Returns/Yields(§1.6).check_return_documentationtreated any non-empty parsedReturns:text as a claim of a real return value, so every void function documented with the generator's own honestReturns:\n None.\nconvention was flagged "has Returns section but doesn't return a value" — agenerate→validateCI pipeline produced spurious warnings on its own freshly generated, correct output. A real generator got the identical false positive, because the check only recognizedast.Return, neverast.Yield/ast.YieldFrom— a generator's genuine yielded value was invisible to it. Fixed: the output-value scan now includesyield <value>, and the literal"None."sentence no longer counts as "claims a return value." On the audit's own dogfood fixture, this dropped the validator'sreturns-category issue count from 5 to 0 with no new categories introduced. - NumPy
Raisessections corrupted the merged docstring, and did so worse afterRaises:generation was added (§1.3). An unrecognized NumPyRaisessection used to be folded into the free-textdescription, producing malformed output (no longer valid Google or NumPy style — it lost its------underline without gaining a:, and floated aboveArgs:/Returns:). Once realRaises:generation existed (see Tier 2a below), regenerating a file with this shape produced two disagreeing Raises blocks: the stale, malformed leftover, and a correct, freshly-generated one naming the same exception.docstring_parser.py's NumPyRaiseshandling now discards the body entirely, matching the design already used for the Google-styleRaises:header: the exception class is always recomputed fresh from the function's actualraisestatement, so there's nothing to merge back in. - String forward-reference type annotations (
-> "ClassName") resolved toAnyinstead of the named type (§1.4).type_analyzer.py'sget_annotation_typehandledast.Constantonly for theNoneannotation; any other constant -- i.e. any quoted forward reference, PEP 484's way of naming a type not yet defined at the annotation's point in the source -- fell through to"any". This hit the tool's own public API:PyCodeCommenter.from_string/from_fileboth declare-> "PyCodeCommenter"(the standard fluent-builder self-return pattern) and gotReturns: Any: ...instead ofReturns: PyCodeCommenter: .... Also affects any self-referencing factory/builder method (confirmed onOrderProcessor.empty(cls, ...) -> "OrderProcessor"andCoordinates.distance_to(self, other: "Coordinates")in the audit's own fixture) and resolves correctly nested inside a generic too (e.g.Optional["ClassName"]), since the fix is in the same recursiveget_annotation_typecall every subscript/union branch already goes through. generatewas not idempotent: regenerating on already-patched output compounded without bound. Every rerun on a file with a defaulted parameter appended another copy of" (default: ...)"onto that parameter's Args: line (rate (float): float value. (default: 0.1)→... (default: 0.1). (default: 0.1)→... (default: 0.1). (default: 0.1). (default: 0.1), unbounded), and a niladic function'sReturns: None.becameReturns: None: None.on the very next regeneration. Root cause:DocstringParserre-parses the tool's own previously generated Args-line suffix and bare"None."return sentence back as if they were free-form author text, andcommenter.pyre-wrapped them instead of recognizing its own prior output — the same defect shape as theRaises:/NumPy-Returns:merge bugs above, for two more tool-generated literals. Pre-existing; reproduces identically on the untouched pre-Tier-2a checkout. Fixed: a new_strip_own_default_annotation(commenter.py) strips any trailing" (default: ...)"suffix from a re-parsed parameter description before deciding what to append — unconditionally, not only when it happens to match the parameter's current default, since the append immediately below always re-adds the correct, current one regardless. (An earlier version of this fix matched only the current value, which left a gap: editing a parameter's default in source betweengenerateruns — a normal workflow — left the stale suffix in place alongside a freshly appended current one, the same compounding failure via a legitimate edit instead of a bare rerun.) The bare"None."return sentence is now recognized and treated as absent (not real preserved text) before the existing-description merge branch runs — falling through correctly to a fresh guess marker if the function has since been edited to actually return a value, rather than keeping the stale text forever. Verified stable across 4 consecutive regeneration passes — including across a changed default value — and on the fulldocstring_generation_fixture.py.- Generation injected an unwanted
TODO(pycodecommenter): describeparagraph into docstrings that were already complete.commenter.py's_generate_function_docstring/_generate_class_docstringtreated "no free-text description paragraph was parsed" as "a description is missing," even when a docstring's summary +Args:+Returns:was already complete by design (a normal, common Google-style shape with no separate prose paragraph). Merging intoslugify's own already-correct docstring — used as the maintainer's own "leave this alone" regression fixture — reproducibly injected the placeholder instead of leaving the docstring untouched. Fixed: the description slot is now only ever filled with real parsed text; it's left empty otherwise, whether or not a prior docstring existed. (A fresh, never-documented function's summary line is itself just derived from the function's name, so a second "nothing to add" paragraph immediately below it added no information beyond what the summary already didn't have — this now no longer appears either. Per-field placeholders —Args:/Returns:/Attributes:/Methods:entries with no other source of truth — are unaffected.) - A hand-wrapped summary sentence spanning multiple physical lines got
split at the wrap point.
DocstringParser.parse()took only the docstring's first physical line as "the summary," so any existing docstring whose summary sentence wrapped across two or more lines (common style throughout this project's own source, e.g.param_utils.py's_walk_until) had its second line treated as a separate description paragraph — severing the sentence mid-thought. Fixed: the summary is now the whole first paragraph (every physical line up to the first blank line or a recognized section header), joined back into one line.
Added¶
- Class
Attributes:descriptions now go through the sameinfer_description()pipelineArgs:already uses, instead of an unconditionalTODO(pycodecommenter): describefor every attribute.OrderProcessor.customer_idnow reads "Unique identifier for the customer." instead of a guess marker; an attribute with no matching name/type signal (e.g.ConfigError.original, typedOptional[Exception]) still correctly gets the guess marker. - Functions that
raisenow get aRaises:section naming the exception class. Previously the generator only ever wroteArgs:/Returns:/Yields:— even for a function whose body raised, which meant runningvalidateagainst the tool's own generated output flagged its own docstrings ("Function raises exceptions {'ValueError'} but has no Raises section"). The exception class at araise SomeError(...)site is a literal token in the source (a fact); only why/when it's raised is unknowable from the AST, so that part is still an explicitTODO(pycodecommenter): describe when this is raised.marker. A bareraise(re-raise) andraise err(an already-constructed instance) are correctly skipped rather than guessed, since neither lets the exception class be read off the raise site without data-flow analysis.docstring_parser.pyalso now recognizesRaises:as a real Google-style section header (it previously didn't), so re-runninggenerateon already-Raises-documented output doesn't glue the Raises block onto theReturns:text above it — the same corruption shape as the pre-existing NumPy-Raisesbug, now pre-empted for this new section too. inference.py's name-pattern vocabulary widened:id/*_id,name/*_name,key/*_key,index/idx,config/configuration/settings/options,result/outputnow produce a real, specific description (e.g.customer_id→ "Unique identifier for the customer.") instead of falling through to the generic type-only fallback or the guess marker. Existing, more specific rules (*_path,is_/has_/can_,count/num) keep priority — the new rules are checked last and never shadow them.
[2.5.0] - 2026-09-05¶
Added¶
--versionflag on the CLI, printing the installed version and exiting 0.
Fixed¶
validate's exception/return checks misattributed a nested function'sraise/returnto the outer function.check_return_documentationandcheck_exception_documentationused a rawast.walk(func_node), which descends into nesteddefbodies — a closure or helper function's ownraise/returnproduced a false-positive WARNING on the enclosing function.commenter.pyalready had the fix for the identical bug class (_walk_function_body, used by_is_generator/_get_return_type) but it was never shared withvalidator.py. Both now shareparam_utils.walk_own_scope._get_class_attributeshad the same class of bug, twice. Aself.x = ...assignment inside a class nested within__init__was misattributed to the outer class'sAttributes:section (fixed with a newparam_utils.walk_skipping_nested_classes, which — unlikewalk_own_scope— still descends into nested functions/closures, since those legitimately share the enclosing__init__'s ownself). Separately,__init__'s own parameters were read directly fromfunc_node.args.argsinstead of the sharedget_all_parameters()/exclude_self_cls()primitive, so keyword-only params and**kwargsgot(any)in theAttributes:section instead of their real type — inconsistent with the correct type already shown for the same parameter in theArgs:section a few lines below.- Negative-number default values (e.g.
x=-1) rendered as(default: unknown).astrepresents a signed numeric literal asUnaryOp(USub, Constant(...)), not a singleConstant—_get_default_valueonly handled the latter. pycodecommenter coverage <file>crashed with a raw traceback (instead of a clean error and exit 1) when the target file couldn't be parsed or didn't exist.CoverageAnalyzer.analyze_file()had no error handling around its ownopen()/ast.parse(), unlikeanalyze_directory()'s per-file wrapping;generate/validatealready handled this case gracefully for a single file.
Removed¶
parameter_descriptions.py, a hardcoded dictionary of parameter descriptions keyed to ~17 exact function names (calculate_area,send_email,connect_to_database, etc.), consulted ahead of the general rule-based inference ininference.py. It never generalized beyond those exact names and fell through safely regardless — removed in favor ofinference.pyalone. A function whose name and parameter exactly matched one of those entries will now getTODO(pycodecommenter): describefor that parameter instead of the old canned sentence, same as any other parameterinference.pyhas no signal for.
[2.4.0] - 2026-08-24¶
Added¶
- New shared
PyCodeCommenter/param_utils.pyprimitive (get_all_parameters/exclude_self_cls), fixing five independent blind spots where positional-only params, keyword-only params,*args/**kwargs, and dataclass/self-assigned class attributes were silently invisible to both docstring generation and validation — all three sites that inspected a function's parameters (the generator's Args-writer, and two validator checks) previously rebuilt that list fromfunc_node.args.argsalone. - Generator functions (containing a
yield) now get aYields:section instead of an incorrectReturns: None. --badge-output PATHflag onpycodecommenter coverage(and the underlyingcoverage.shields_badge_dict()), emitting a shields.io endpoint-badge JSON file for the coverage percentage.- CI now runs the full test suite, including
inference.py's doctest examples, on every push/PR againstmain(.github/workflows/tests.yml, new; matrix over Python 3.9-3.12).
Changed¶
- README repositioned to be explicit that generated prose the tool can't extract from the AST is a marked placeholder needing review, not finished documentation, and to name related tools (pydoclint, interrogate) instead of implying PyCodeCommenter is alone in this space. The Roadmap's AI-generation item was replaced with a struck-through line plus a warning block spelling out the draft-and-review-only constraint, so it can't be misread as an open, unclaimed feature.
Fixed¶
- Guessed (as opposed to AST-extracted or preserved) docstring content is now
marked with
GUESS_MARKER = "TODO(pycodecommenter): describe"instead of being presented as finished prose. This string is already invalidator.py's placeholder blacklist, so generated output with unresolved guesses now fails the tool's own quality check instead of silently passing it. _get_class_attributesnow also picks upself.x = ...assignments anywhere in__init__'s body (not just__init__'s own parameters) and class-levelAnnAssignfields (covers@dataclass-style classes with no__init__written in source).@classmethod-decorated functions no longer documentclsin the generated Args section (the generator only excludedself; the validator already excluded both).- Python 3.9 import crash in
config.py. A live (non-deferred)Exception | Nonedefault-argument annotation needs Python >=3.10 —type.__or__for builtin/exception types doesn't exist on 3.9, so evaluating it at class-definition time (import time) raisedTypeError: unsupported operand type(s) for |: 'type' and 'NoneType'. Sincecli.pyimportsconfig.py, this broke collection for the entire test suite on 3.9. Pre-existing since v2.1.0 (2026-06-10), roughly three months before this release — caught by this release's own new CI matrix (see above) on its first real run. Fixed withOptional[Exception]plusfrom __future__ import annotations, matching the guardinference.pyalready uses for its ownstr | Noneusage.
[2.3.0] - 2026-08-22¶
Follow-up work from an internal engineering audit (docs/dev-notes/vulnerabilities.md),
plus a documentation-fidelity pass (Phase 6) found during pre-release review.
Breaking Changes¶
- Minimum Python version raised from 3.8 to 3.9.
pyproject.toml'srequires-pythonis now>=3.9; theProgramming Language :: Python :: 3.8classifier was removed. This was forced by the newlibcstdependency (see below), whose current release requires Python >=3.9. Any user still on Python 3.8 will no longer be able to install new releases of this package.
Added¶
- New runtime dependency:
libcst>=1.1.get_patched_code()(the code that inserts/updates generated docstrings) was rewritten to apply edits via libcst's concrete syntax tree instead of line-number arithmetic on raw source text. This is a hard dependency, not optional — installingpycodecommenternow also installslibcst. generateandvalidateCLI commands now accept a directory as well as a single file, recursively collecting.pyfiles (mirroringcoverage's existing directory support).--fail-below THRESHOLDflag on thecoverageCLI command; exits 1 when coverage is below the threshold..pycodecommenter.yamlconfig is now actually read by the CLI: its top-levelexcludelist defaults-e/--excludeon all three subcommands, andcoverage.thresholddefaults--fail-below. Both remain overridable by explicit CLI flags. (Previouslyconfig.pyexisted but nothing called it.)- NumPy-style docstrings (
Parameters/Returns/Raiseswith dash-underlined headers) are now parsed as input, alongside Google and Sphinx style. Sphinx-style input also gained:type name: TYPEsupport.
Fixed¶
- One-line function/class definitions could corrupt the file when patched.
def foo(): return 1(body on the same physical line as thedef) would have its generated docstring inserted before thedefline instead of inside the function, producing aSyntaxError. Every one-liner shape reproduced this:pass,return, multiple;-separated statements,async def, and one-lineclassbodies. Fixed by the libcst rewrite above, which converts a one-liner body to a proper indented block before inserting. (This is a different, narrower bug than the "multi-line signature corruption" originally suspected from the audit — see Notes below.) PyCodeCommenter.validate()now passes the real file path to the validator instead of a hardcodedNone, so validation report locations show the actual file (e.g.src/api.py:12:my_func) instead ofcode:12:my_func.PyCodeCommenter.check_coverage()now uses the real file path instead of the leaked"<string>"placeholder.docstring_parser.py: multi-lineArgs:parameter descriptions in an existing docstring being re-parsed were silently truncated unless the continuation line was indented by exactly 8 spaces. Any other indentation (including 4 spaces — the same indent this project's own generator uses) lost the continuation text on merge. Now any non-blank continuation line is recognized regardless of indentation depth.generate/validatedirectory mode's default exclude list (__pycache__,.git,.venv,venv,env,.eggs) missed common vendor/build directories and used raw substring matching, causing two real problems:environment_config.pywas silently skipped ("env"is a substring of"environment"), and files inside.tox/.../site-packages/were not skipped — withgenerate --inplace, writing generated docstrings into vendored third-party source. Expanded the default list (added.tox,.nox,__pypackages__,site-packages,build,dist,.egg-info,.mypy_cache,.pytest_cache,node_modules) and switched to exact-path-component matching.coverage's directory mode had the same two problems independently (a separate, duplicated exclude list) and is fixed the same way, plus a related bug: passing any custom--excludepattern previously replaced its default list entirely rather than adding to it, silently losing.venv/.gitprotection.- A parameter's documented type (e.g.
x (int):) was silently downgraded tox (any):on regeneration whenever static type inference had nothing to work with (no annotation on the parameter) — the type was parsed out of the existing docstring but never stored or consulted. A real static annotation still always wins; the docstring-parsed type is now used as a fallback instead of being discarded. - An existing NumPy-style docstring was not recognized as such, so its
entire body was treated as free-text and a fresh, auto-generated
Google-style
Args:/Returns:section was appended underneath it — documenting the same parameter twice, in two styles, in one docstring. - Generated filler text for dunder methods (e.g.
__init__) rendered with broken spacing —"Original of the init ."— becausename.replace('_', ' ')turns every underscore into a space, including the leading/trailing pair(s) dunder names have.
Changed¶
ValidationStats.infosrenamed to.info(matches the existing"info"key in JSON/markdown output)..infosremains available as a backward-compatible property alias.
Notes¶
- Two "CRITICAL" bugs originally suspected from the audit — corruption on multi-line function signatures, and a line-shift bug when patching multiple functions in one file — were investigated and do not reproduce against this version or the pre-libcst version; both were verified directly against the audit's own examples plus additional stress tests. No code changes were made for either.
- A third suspected bug —
ast.walk()double-counting nested functions/classes invalidate_all()'s stats — also does not reproduce;ast.walk()visits every node exactly once regardless of nesting depth.validator.py's counting logic is unchanged. - PEP 604 union rendering (
int | strcurrently renders asUnion[int, str]in generated docstrings, pertype_analyzer.py) was identified as a real, minor issue but deliberately deferred: fixing it breakstest_type_analyzer.py::test_union_annotation, which asserts the oldUnion[...]output, and no test-file changes were in scope for that part of the work.
[2.2.0] - 2026-07-12¶
Added¶
- Decorator-aware validation —
@propertysetter and deleter variants no longer produce false-positive "missing Returns section" warnings.@classmethod(cls) and@staticmethodwere verified already correct. --output-format jsonflag on bothvalidateandcoveragesubcommands. Prints machine-readable JSON to stdout; text output is completely unchanged.ValidationReport.to_dict()updated to spec-compliant shape:stats.total,stats.info,stats.coverage_percentage; issues now includeline(integer),severity(uppercase),check,message._has_raises_section()helper invalidator.py— recognisesRaises:(Google),Raises\n(bare header), and:raises(Sphinx, trailing space prevents false matches).llms.txtfile at repo root — AI agent and LLM crawler discovery file.- SEO/AEO optimisation — keyword-rich README, PyPI classifiers/keywords, per-page meta tags on all
doc pages, JSON-LD
SoftwareApplicationschema on docs homepage.
Changed¶
test_validation.pyrewritten from script-style to proper pytest (17 tests, all pass).pyproject.tomlkeywords expanded to 20 terms; classifiers expanded withCode Generators,Libraries :: Python Modules,Environment :: Console,Operating System :: OS Independent.mkdocs.ymlenriched withsite_author, explicitlanguage: en,navigation.indexes,search.share,toc.follow, andmetamarkdown extension for per-page SEO tags.
[2.0.0] - 2026-01-25¶
Major Release - Complete Rewrite¶
Added¶
- Comprehensive Validation System - 6 types of documentation checks
- Coverage Analysis - Project-wide documentation metrics
- Modern Type Support - PEP 604 unions, PEP 585 generics
- Async Function Support - Full support for
async def - Multiple Export Formats - JSON, Markdown, console output
- Smart Docstring Parsing - Preserves existing documentation
- Structured AST Traversal - NodeVisitor pattern for reliability
- Professional Reporting - Actionable error messages with suggestions
- CI/CD Ready - Easy integration with pipelines
Changed¶
- Replaced basic type inference with comprehensive
TypeAnalyzer - Improved docstring generation with better templates
- Enhanced error handling with proper logging
- Better handling of edge cases and malformed code
Fixed¶
- Duplicate
_infer_typemethods consolidated - Brittle patching logic made robust
- Import resolution issues
- Unicode/encoding handling
Breaking Changes¶
- Minimum Python version: 3.8+
- Some internal APIs changed (public API remains compatible)
[1.0.0] - Earlier Version¶
- Basic docstring generation
- Template-based descriptions
- File and string input support