I’m building a desktop GenAI code assistant that should:
Read the user’s entire project folder(potentially thousands of files).
Provide context-aware code suggestions and debugging help.
Support multiple LLM providers(OpenAI, Gemini, DeepSeek) using API keys.
Sending the entire codebase in each API request is impractical because:
LLM context windows are limited.
Huge latency and network cost if we send megabytes of code each request.
Cost scales linearly with tokens.
I want to understand how real-world tools like GitHub Copilot and Cursor handle this:
How do they provide accurate code suggestions without sending the whole repository every request?