feat: Implement smart time-based rate limiting v0.1.8

🎯 Major Rate Limiting Improvement - Research-Friendly Design

New Rate Limiting System:
✅ 12 calls per 5-minute window (vs 8 per session)
✅ 8 calls per 30-second burst protection
✅ Automatic time-based reset (no LM Studio restarts needed)
✅ Clear feedback with remaining calls and reset times

Key Benefits:
- Supports extended research sessions without interruption
- Prevents LLM spam while allowing legitimate research workflows
- Users can work continuously without restarting LM Studio
- Intelligent burst protection prevents overwhelming websites

Technical Implementation:
- Sliding window algorithm with timestamp tracking
- Dual-layer protection: burst + window limits
- Automatic cleanup of expired call history
- User-friendly error messages with precise wait times

This addresses the core user feedback that session-based limits were too restrictive for normal research use cases while maintaining responsible web scraping practices.

Breaking Change: Rate limiting behavior changed from session-based to time-based
Migration: No action needed - new system is more permissive
This commit is contained in:
Jay
2025-08-10 23:47:42 -05:00
parent d380736ea0
commit e1f6fff3fe
4 changed files with 65 additions and 14 deletions
+20 -1
View File
@@ -1,4 +1,4 @@
# 🌐 WebFetch.MCP v0.1.7
# 🌐 WebFetch.MCP v0.1.8
**Live Web Access for Your Local AI — Tunable Search & Clean Content Extraction**
@@ -119,6 +119,25 @@ In LM Studio:
| DEBUG | false | Debug logging |
| DETAILED_LOG | true | Detailed log output |
# ⏱️ Smart Rate Limiting
WebFetch.MCP uses intelligent time-based rate limiting designed for real research workflows:
### **📊 Rate Limits:**
- **12 calls per 5-minute window** - Generous limit for research sessions
- **8 calls per 30-second burst** - Prevents LLM spam while allowing quick queries
- **Automatic reset** - No need to restart LM Studio between research sessions
### **🎯 Why This Works Better:**
- ✅ **Research-friendly** - Supports extended research sessions
- ✅ **Anti-spam protection** - Prevents runaway LLM tool calling
- ✅ **No restarts needed** - Limits reset automatically over time
- ✅ **Clear feedback** - Shows remaining calls and reset times
### **📈 Example Usage Patterns:**
- **Quick research**: 5-8 rapid calls, then brief pause
- **Extended research**: 12 calls spread over 5 minutes
- **Continuous work**: Limits reset as you work, no interruption
# 📊 Example Usage
**Search**
```
+1 -1
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env node
/**
* WebFetch.MCP - Log Monitor Utility v0.1.7
* WebFetch.MCP - Log Monitor Utility v0.1.8
* Real-time log monitoring for WebFetch.MCP server
*
* This utility monitors the detailed log file generated by the WebFetch.MCP server
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "webfetch-mcp",
"version": "0.1.7",
"version": "0.1.8",
"type": "module",
"private": false,
"description": "Production-ready MCP server for web search and content extraction. Live web access for your local AI with tunable search and clean content extraction.",
+43 -11
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env node
/**
* WebFetch.MCP v0.1.7
* WebFetch.MCP v0.1.8
* Live Web Access for Your Local AI — Tunable Search & Clean Content Extraction
*
* A production-ready Model Context Protocol (MCP) server that provides web search
@@ -16,7 +16,7 @@
*
* @author Jay Leon (@manull)
* @license MIT
* @version 0.1.7
* @version 0.1.8
* @repository https://github.com/manull/webfetch-mcp
*
* Copyright (c) 2025 Jay Leon (@manull)
@@ -43,26 +43,58 @@ const SEARXNG_BASE = process.env.SEARXNG_BASE || "http://localhost:8080";
const DEBUG = process.env.DEBUG === "true";
const DETAILED_LOG = process.env.DETAILED_LOG !== "false"; // Default to true
// Simple call tracking
let callCount = 0;
const MAX_CALLS = 8;
const startTime = Date.now();
// Time-based rate limiting - more user-friendly approach
const RATE_LIMIT_WINDOW_MS = 5 * 60 * 1000; // 5 minutes
const MAX_CALLS_PER_WINDOW = 12; // Allow more calls but over time
const BURST_LIMIT = 8; // Max calls in quick succession
const BURST_WINDOW_MS = 30 * 1000; // 30 seconds
let callHistory = []; // Array of timestamps
const checkCallLimit = () => {
callCount++;
const remaining = MAX_CALLS - callCount;
const now = Date.now();
if (callCount > MAX_CALLS) {
// Clean up old calls outside the main window
callHistory = callHistory.filter(timestamp => now - timestamp < RATE_LIMIT_WINDOW_MS);
// Check burst limit (quick succession)
const recentCalls = callHistory.filter(timestamp => now - timestamp < BURST_WINDOW_MS);
if (recentCalls.length >= BURST_LIMIT) {
return {
limited: true,
message: `🛑 **Rate Limit Reached**: You've made ${callCount} tool calls. Please restart LM Studio to reset the limit, or try to work with the information already gathered. Consider being more specific in your queries to get better results with fewer calls.`
message: `🛑 **Burst Limit Reached**: ${BURST_LIMIT} calls in ${BURST_WINDOW_MS/1000} seconds. Please wait ${Math.ceil((BURST_WINDOW_MS - (now - recentCalls[0]))/1000)} seconds before making more requests. This prevents overwhelming websites and ensures reliable service.`
};
}
// Check overall window limit
if (callHistory.length >= MAX_CALLS_PER_WINDOW) {
const oldestCall = Math.min(...callHistory);
const resetTime = Math.ceil((RATE_LIMIT_WINDOW_MS - (now - oldestCall)) / 1000 / 60);
return {
limited: true,
message: `🛑 **Rate Limit Reached**: ${MAX_CALLS_PER_WINDOW} calls in ${RATE_LIMIT_WINDOW_MS/1000/60} minutes. Please wait ${resetTime} minute(s) for the limit to reset. This ensures responsible web scraping and prevents server overload.`
};
}
// Add current call to history
callHistory.push(now);
// Provide helpful warnings
const remaining = MAX_CALLS_PER_WINDOW - callHistory.length;
const recentCount = recentCalls.length + 1; // +1 for current call
if (remaining <= 2) {
return {
limited: false,
warning: `⚠️ **${remaining} calls remaining** - Please use them wisely.`
warning: `⚠️ **${remaining} calls remaining** in this ${RATE_LIMIT_WINDOW_MS/1000/60}-minute window.`
};
}
if (recentCount >= BURST_LIMIT - 2) {
return {
limited: false,
warning: `⚠️ **${BURST_LIMIT - recentCount} quick calls remaining** - Consider spacing out requests.`
};
}