feat: add SerpBase (Google) search engine via REST API - #323
gefsikatsinelou wants to merge 1 commit into
Conversation
WebSearchAgent now reads SERPBASE_API_KEY and queries the SerpBase Google Search API when set, falling back to DuckDuckGo (DDGS) when the key is missing or the API call fails. No new dependencies; reuses requests. Docs updated in docs/agents/acquisition/web-search.md.
mikegros
left a comment
There was a problem hiding this comment.
This is great and having approaches and fallbacks are a great idea. I need to test it out more, but wanted to bring up 2 things:
-
There was some stuff removed and I am not clear why (see my comment).
-
It would be really great to also have Tavily as an optional search if there is an API key set for it. So that for the user, the precedence went Serp -> Tavily -> DDGS. This would be nice from the perspective of giving the user the ability to use different search backends without affecting current behavior.
I could always implement (2) in the future if you would rather just get this merged, but I'd like clarity on (1) first.
| "duckduckgo-search (DDGS) is required for WebSearchAgentGeneric." | ||
| ) | ||
|
|
||
| def _id(self, hit_or_item: dict[str, Any]) -> str: |
There was a problem hiding this comment.
Why are _id and _citation being removed here? I think they are used in _materialize? Is there a reason they are removed?
Summary
Adds an optional Google search backend to
WebSearchAgentvia the SerpBase REST API, with a graceful fallback to the existing DuckDuckGo (DDGS) path. No new dependencies.Background
WebSearchAgent._searchcurrently relies solely onddgs(DuckDuckGo). In practice DuckDuckGo rate-limits aggressively for agent-style bursts (and is blocked from some datacenter IPs), so research runs that need Google coverage have no option today. SerpBase is a lightweight Google Search Results API (structuredorganic_resultsJSON, no scraping, no headless browser) that slots into the existing_search→_materializeflow without touching the acquisition graph.Changes
src/ursa/agents/acquisition_agents.pyWebSearchAgent.__init__readsSERPBASE_API_KEYfrom the environment._serpbase_search()— GEThttps://api.serpbase.dev/google/search?q=...&num=<max_results>, mapsorganic_results→ the same{title, href, body, position}dict shape DDGS produces (so_id,_materialize, and_citationwork unchanged)._search()prefers SerpBase when the key is set; falls back to DDGS when the key is missing, the API errors, or returns no results.logger(used by the fallback warning); no behavioral change elsewhere.docs/agents/acquisition/web-search.md— documents the new env var, the fallback behavior, and that no new dependency is required.Design decisions
SERPBASE_API_KEYvia env varUNPAYWALL_EMAIL,URSA_TEXT_EXTENSIONSare read fromos.environ).requestsacquisition_agents.py; no new dependency._id/_materialize/_citationoperate onhref/title/body— zero changes needed downstream.Testing
python -m py_compile src/ursa/agents/acquisition_agents.pypasses.SERPBASE_API_KEYset:_searchreturns Googleorganic_resultsmapped to{title, href, body, position}._searchbehaves exactly as before (DDGS only).