fix(ml): resolve Finding and dict type mismatch in deduplication #259 - #346
Open
prasiddhi-105 wants to merge 4 commits into
Open
fix(ml): resolve Finding and dict type mismatch in deduplication #259#346prasiddhi-105 wants to merge 4 commits into
prasiddhi-105 wants to merge 4 commits into
Conversation
|
🎉 Thank you @prasiddhi-105 for submitting a Pull Request! We're excited to review your contribution. ✅ Before Review
⚡ Want faster reviews and contributor support? Join our Discord community: 🔗 https://discord.gg/FcXuyw2Rs Maintainers and mentors are active there and can help resolve blockers quickly. Happy Contributing! 🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Linked issue
Closes #259
What this PR does
Fixes a type annotation and attribute access mismatch in the deduplication pipeline where
dictinputs could cause silent empty string extraction inembed_findings(). Also widens exception handling duringSentenceTransformerinitialization inembedder.pyto prevent crashes caused by network errors, GPU memory exhaustion, or model load failures.Type of change
ML tier (if applicable)
Stack affected
Changes
Backend
backend/app/ml/embedder.py:_extract_text()to safely handle both PydanticFindingobjects and raw Pythondictinstances.SentenceTransformerloading to catch all errors (e.g.,OSError,ConnectionError,OutOfMemoryError) and fail gracefully.embed_findings()type hints toList[Union[Finding, Dict[str, Any]]].backend/app/ml/deduplicator.py:deduplicate()and internal wrapperembed_findings()to explicitly allowUnion[Finding, Dict[str, Any]].Frontend
New dependencies
Database / schema changes
Testing
How did you test this?
python -m pytest(109 passing tests)._extract_textreturns non-empty strings for bothFindingmodel instances and standard dictionary inputs.Checklist
console.erroror unhandled Python exceptions introducedrequirements.txt/package.jsonupdated if new dependencies added.pkl,.pt, etc.) are gitignored, not committedAnything reviewers should focus on
Please check the dual
Finding/dicthandling in_extract_text()and the broadenedExceptioncatch duringSentenceTransformerinitialization.Screenshots (if UI changed)
N/A (Backend logic fix)