Repository navigation
Add similar-problem search and fix vector database pipeline - #6
Merged
Merged
Conversation
- Add autosetter.similarity: after extraction, embed problem.json and look up similar stored problems in Qdrant (optional, --no-similarity) - Add vector_database/search.py (ProblemSearcher) with deterministic Codeforces URL resolution; test_search.py uses it instead of Gemini - Fix embed/upload resume: track lines consumed and stop on a failed batch so the next run retries it instead of skipping it - Use EMBEDDING_BATCH_SIZE in the embedder - Keep missing difficulty as None instead of the string "None" - Drop per-problem sleep in the Codeforces API collector - Fix similarity test stub to match the current extraction prompt Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
problem.json, AutoSetter embeds it and looks up the closest stored problems in Qdrant. Matches are saved tosimilar_problems.json. The step is optional (--no-similarity,AUTOSETTER_SIMILARITY=0) and is skipped with a warning if the dependencies or Qdrant are missing.vector_database/search.py: newProblemSearcherused by both the CLI andscripts/test_search.py. Codeforces URLs are built from the problem ID (including gym contests), replacing the Gemini call intest_search.py.pipeline.py: the resume position now counts lines read fromproblems.jsonl. A failed upload batch stops the run so the next run retries it. Before, a later successful batch moved the position past the failed one, and it was never uploaded.EMBEDDING_BATCH_SIZEis now used by the embedder.Noneinstead of the string"None".Testing
pytest: 81 passed.scripts/test_search.py "shortest path in weighted graph…"returns Dijkstra? (0.85). A reworded Theatre Square problem run throughfind_similar_problemsreturns Theatre Square (0.85).Notes
vector_database/requirements.txtin the same environment asautosetterand have Qdrant running onlocalhost:6333.🤖 Generated with Claude Code