Small, dependency‑light Python utilities for any single‑file HTML app with large embedded JavaScript/CSS. Analyze structure, split inline code into external assets, and clean/pretty‑print — without brittle regex hand‑parsing.
Handy for code review, refactoring legacy single‑file pages, diffing, and shrinking the context you paste into an LLM (extract the JS instead of the whole HTML).
js_analyzer.py— primary tool: analyze · extract · cleanhtml_cleaner.py/separator.py— simpler single‑purpose variants
Python 3.9+ and:
pip install beautifulsoup4 esprima jsbeautifier # jsbeautifier only needed for --cleanParsing uses the
esprimapackage (≈ES2017). Very new syntax may not parse.
| Goal | Command |
|---|---|
| Structural report (functions, variables, classes, scope leaks, global mutations) | js_analyzer.py input_file [-n -g -u -a -j] |
Split inline JS/HTML into *_extracted.js + *_extracted.html |
js_analyzer.py input_file -e |
| Clean & pretty‑print inline JS + CSS | js_analyzer.py input_file --clean |
js_analyzer.py is the maintained, CLI‑driven tool; html_cleaner.py and
separator.py are simpler variants kept for convenience.
AST‑based structural analysis of JavaScript embedded in HTML, with safe extraction.
js_analyzer.py [-h] [-n] [-a] [-u] [-j] [-f] [-v] [-c] [-g] [-e] [--clean] input_file
-n Prefix results with original-HTML line numbers
-a Include local variables (not just globals)
-u List unintended global "leak" assignments
-g Track global-variable mutations inside scopes
-f | -v | -c Show only functions | variables | classes
-j Emit results as JSON
-e Extract JS/HTML to <input_file>_extracted.js and <input_file>_extracted.html
--clean Clean & pretty-print to <input_file>_clean.html FIRST, then run the report/-e on it
Examples
js_analyzer.py page.html -n -g -u # human-readable structural report
js_analyzer.py page.html -j > report.json # machine-readable
js_analyzer.py page.html -e # -> page_extracted.js + page_extracted.html
js_analyzer.py page.html --clean # -> page_clean.html
js_analyzer.py page.html --clean -e # -> page_clean.html, then page_clean_extracted.js/.htmlWith
--clean, cleaning happens first and all remaining steps use the cleaned file as the input — so the report's line numbers and any-eoutput (<input_file>_clean_extracted.*) refer to<input_file>_clean.html.
Notes & limitations
-econcatenates multiple inline<script>blocks; duplicate top-levellet/constacross blocks can collide (you'll be warned before proceeding).- HTML re-serialization can normalize whitespace and collapse delicate inline spans.
js_analyzer.pyincludes a post-serialization repair pass (repair_title_layout()) that you can adapt in the source to keep such markup byte-for-byte.
html_cleaner.py input_file— clean mixed tab/space indentation and tidy inline JS/CSS, writing<input_file>_clean.html. (Equivalent tojs_analyzer.py input_file --clean.)separator.py input_file— a simpler, standalone splitter: writes<input_file>_extracted.html+<input_file>_extracted.jsand a<input_file>_report.txtreport that lists only global variables and function declarations. For richer analysis (classes, methods, local scope, leaks, mutations) usejs_analyzer.py.
GPL‑3.0