feat(search): add SQLite FTS5 post-build search index script - #801
jeevantelukula wants to merge 1 commit into
Conversation
|
What are we actually going to use these databases for? Are they just going to be distributed alongside the documentation or is this expected to hook into the existing search function? If it's the latter, couldn't this be achieved with a custom scorer function. |
It is going to be distributed alongside the documentation. The existing search functionality which uses searchindex.js is good enough for manual search filtering. Ref: https://github.com/jeevantelukula/ti-processor-sdk-linux-skills/tree/master/mcp |
|
In that case I would prefer if the packaging / release preparation scripts specifically for ti.com be carried somewhere internally. I don't really want to include this in the public builds as it increases build time and results in artifacts too big to distribute here. |
Sphinx's searchindex.js stores only boolean term membership with no TF-IDF or field weighting, works good for manual browser search but gives poor ranking for programmatic (MCP/AI agent) search queries. Add build_searchindex.py that optionally runs after sphinx-build via html make target, parses HTML output, extracts the weighted fields (based on title, code_blocks, image_alt and body_text) and writes searchindex.db with a BM25-ranked FTS5 index. MCP server queries this DB to return accurately ranked doc pages to AI agents. Also add beautifulsoup4 additionally to docker deps. Signed-off-by: Telukula Jeevan Kumar Sahu <j-sahu@ti.com>
f1479c9 to
0f67d0c
Compare
Good point. Although, the addtional size we are adding with this db is around 2-3MB for a device. Considering the number of devices we have, I have made this optional. So, github pages won't generate this by default and internally using |
Sphinx's searchindex.js stores only boolean term membership with no TF-IDF or field weighting, works good for manual browser search but gives poor ranking for programmatic (MCP/AI agent) search queries.
Add build_searchindex.py that runs after sphinx-build via html make target, parses HTML output, extracts the weighted fields (based on title, code_blocks, image_alt and body_text) and writes searchindex.db with a BM25-ranked FTS5 index. MCP server queries this DB to return accurately ranked doc pages to AI agents.
Also add beautifulsoup4 additionally to docker deps.