Skip to content

Add similar-problem search and fix vector database pipeline - #6

Merged
IamShreshth merged 1 commit into
mainfrom
similar-problem-search
Oct 4, 2026
Merged

IamShreshth merged 1 commit into
mainfrom
similar-problem-search

Conversation

@IamShreshth

Copy link
Copy Markdown
Collaborator

Summary

  • Similar-problem search: after extracting problem.json, AutoSetter embeds it and looks up the closest stored problems in Qdrant. Matches are saved to similar_problems.json. The step is optional (--no-similarity, AUTOSETTER_SIMILARITY=0) and is skipped with a warning if the dependencies or Qdrant are missing.
  • vector_database/search.py: new ProblemSearcher used by both the CLI and scripts/test_search.py. Codeforces URLs are built from the problem ID (including gym contests), replacing the Gemini call in test_search.py.
  • Resume fix in pipeline.py: the resume position now counts lines read from problems.jsonl. A failed upload batch stops the run so the next run retries it. Before, a later successful batch moved the position past the failed one, and it was never uploaded.
  • Smaller fixes:
    • EMBEDDING_BATCH_SIZE is now used by the embedder.
    • A missing difficulty is stored as None instead of the string "None".
    • Removed the 0.1s sleep per problem in the Codeforces API collector.
    • Fixed the similarity test stub, which matched text no longer in the extraction prompt.

Testing

  • pytest: 81 passed.
  • Live Qdrant with the 12 sample problems: scripts/test_search.py "shortest path in weighted graph…" returns Dijkstra? (0.85). A reworded Theatre Square problem run through find_similar_problems returns Theatre Square (0.85).
  • Resume: forced a batch to fail. Run 1 uploaded 4 of 12 problems and stopped; run 2 uploaded the other 8, with no duplicates.

Notes

  • To use the similarity check, install vector_database/requirements.txt in the same environment as autosetter and have Qdrant running on localhost:6333.
  • The database currently holds only the 12 sample problems. It needs a real Codeforces dataset with statements before the matches are useful.

🤖 Generated with Claude Code

- Add autosetter.similarity: after extraction, embed problem.json and
  look up similar stored problems in Qdrant (optional, --no-similarity)
- Add vector_database/search.py (ProblemSearcher) with deterministic
  Codeforces URL resolution; test_search.py uses it instead of Gemini
- Fix embed/upload resume: track lines consumed and stop on a failed
  batch so the next run retries it instead of skipping it
- Use EMBEDDING_BATCH_SIZE in the embedder
- Keep missing difficulty as None instead of the string "None"
- Drop per-problem sleep in the Codeforces API collector
- Fix similarity test stub to match the current extraction prompt

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@IamShreshth
IamShreshth merged commit 0c1038f into main Oct 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant