Helix Ecosystem · CIS · CAP · CISS · CIB
External perception and progressive information sniffing microservice for the Helix ecosystem.
Tentacle is the "world perception organ" of the Helix digital lifeform. It implements progressive information foraging based on information foraging theory, enabling two-phase web content extraction:
- Low-resolution scan: Generate document topography with keyword hit density
- High-resolution extract: Pull raw text only from high-value sections
- Agent-Programmable Perception: Full
KeywordFiltersupport (include,exclude,boost) applied consistently across search, scan, and extract. - Domain & Site Configuration: Load domain-specific defaults and restrict searches to specific sites via
--domainand--site. - Cookie/Session Support: Authenticated scraping using Netscape-format cookie files.
- Content Quality Scoring: Automatic quality estimation for each DOM section (purity, position, tag diversity).
- Multi-level Filtering:
--filter-level(none,standard,strict) to control aggressiveness of content filtering.
Tentacle supports two running modes sharing the same core engine:
Run as a REST API microservice within the Helix ecosystem.
- Used by Anaphase (Helix's prefrontal cortex)
- Full HXR audit logging, Trace ID propagation
- Integrated with Tuck gateway for security
- Disables CLI interface
Run as a standalone CLI tool for local development and usage.
- Human-friendly CLI interface
- Local logging, basic SSRF protection
- Independent of Helix deployment
# Install from source (editable)
pip install -e .
# Or install with dev dependencies
pip install -e ".[dev]"First enable standalone mode:
export TENTACLE_MODE=standalone# Basic scan
tentacle scan https://example.com --keywords "AI,Agents" --format rich
# Scan with advanced filtering
tentacle scan https://example.com \
--keywords "AI" \
--require "report" \
--exclude "advertisement" \
--boost "2025:2.0" \
--filter-level strict \
--format table# Extract raw text from specific sections
tentacle extract https://example.com --sections sec_001,sec_003# Search the web
tentacle search "LLM reasoning" --limit 3
# Search only within specific sites
tentacle search "AI anxiety" --site "zhihu.com,bbc.com" --limit 2
# Search using a domain configuration (loads default keywords/sites)
tentacle search "electronics" --domain trade --limit 5# Place cookie file in ./cookies/zhihu.txt
tentacle scan https://www.zhihu.com/people/me --cookie zhihu.txt --format table# Explore document interactively
tentacle explore https://example.comStart the API server:
# Default mode is embedded
uvicorn tentacle.api.main:app --host 0.0.0.0 --port 8021The server will be available at http://localhost:8021
| Method | Path | Description |
|---|---|---|
| POST | /v1/tentacle/scan |
Phase 1: Scan URL to get topography |
| POST | /v1/tentacle/extract |
Phase 2: Extract raw text from sections |
| POST | /v1/tentacle/search |
Web search proxy |
| POST | /v1/tentacle/feedback |
Submit feedback for model evolution |
| GET | /health |
Health check |
| GET | /metrics |
Prometheus metrics (optional) |
Interactive API docs are available at http://localhost:8021/docs
All configuration is done via environment variables or .env file. See .env.example for all available options.
Key configurations:
TENTACLE_MODE: Running mode (embedded/standalone)TENTACLE_PORT: API server port (default: 8021)TENTACLE_SEARCH_PROVIDER: Search engine provider (duckduckgo/serpapi)TENTACLE_DOMAINS_DIR: Path to domain configuration YAML filesTENTACLE_MAX_SNIPPET_SIZE: Maximum snippet size per sectionTENTACLE_FILTER_LEVEL: Global default filter level (none/standard/strict)
- SSRF Protection: Blocks access to private IP addresses by default
- Input Validation: Strict schema validation for all inputs
- Rate Limiting: Handled by Tuck gateway in embedded mode
- Audit Logging: Full request/response logging with trace IDs (HXR compatible)
# Run tests
pytest
# Lint code
ruff check .
# Format code
ruff format .MIT