Dormant since December 2025 (verified 2026-10-02)

Repo BeehiveInnovations/pal-mcp-server (the renamed Zen MCP) is not archived, with 11,762 stars and 1,043 forks, but its last push and latest release (v9.8.2, a security fix for path-traversal handling) are both dated 2025-12-15, so it has had no commits for about nine months. GitHub reports the licence as “Other” (NOASSERTION); the earlier note said Apache-2.0, so check the LICENSE file before reuse. README: enabled-by-default tools are chat, planning, code review, debugging and similar; analyze, refactor, secaudit and other tools are disabled by default; clink bridges Claude Code, Gemini CLI, Codex CLI, Qwen Code CLI and Cursor; providers include OpenAI, Gemini, Azure OpenAI, Anthropic, Grok, Ollama, OpenRouter, DIAL; Python 3.10+, uvx install. Gemini CLI was sunset on 2026-06-18 (replaced by Antigravity CLI), so Gemini CLI mentions below are legacy, and model names such as O3, GPT-5 and Gemini Pro predate the 2026 lineups (see gpt-5-and-gpt-6-family); the “50+ models” count is the README’s claim.

PAL MCP Server

by Beehive Innovations

Your AI’s PAL – a Provider Abstraction Layer for Model Context Protocol

See https://github.com/BeehiveInnovations/pal-mcp-server

Features

  • Multi-Model Orchestration: Connect Claude Code, Gemini CLI, Codex CLI, and IDE extensions to 50+ AI models simultaneously (Gemini, OpenAI, Anthropic, Grok, Azure, Ollama, and more)
  • Conversation Threading: Maintain full context across different AI tools and models - discussions flow seamlessly between CLAUDE, Gemini Pro, O3, GPT-5, and other models
  • CLI Subagents (“clink”): Launch isolated AI CLI instances from within your current session - spawn specialized subagents for code reviews, planning, or debugging without consuming your primary context window
  • Persistent Context: Maintains conversation context even after CLAUDE’s memory resets, enabling truly persistent AI collaboration across multiple sessions
  • Vision Support: Analyze images, diagrams, screenshots with vision-capable models - works seamlessly with all tools and conversation threading
  • Extended Context Windows: Delegate to models with massive context limits (Gemini’s 1M tokens) for analyzing large codebases
  • Local Model Support: Run Llama or Mistral locally via Ollama for complete privacy
  • Smart Token Management: Automatically handles large prompts as files to work around MCP’s ~25K token combined request+response limit
  • Code Analysis Tools: Built-in debugging, security audits, documentation generation, and collaborative planning
  • API Lookup: Access current API information (not training-data-based)

Superpowers

PAL MCP Server transforms isolated AI coding assistants into a coordinated development team. Instead of being limited to a single model, developers can orchestrate responses across different AI systems to gain diverse perspectives on coding challenges.

Who this is for:

  • Developers using Claude Code, Cursor, VS Code extensions, or other MCP-compatible tools
  • Teams wanting to leverage multiple LLMs in a single workflow
  • Engineers working on complex codebases requiring diverse AI perspectives
  • Developers needing persistent context across AI sessions
  • Privacy-conscious developers who want local model options

What you gain:

  • Multi-Model Collaboration: Coordinate Gemini Pro, O3, GPT-5, and 50+ other models to get the best analysis for each task
  • Context Revival: Even when one model’s context resets, others can remind it of previous discussions
  • Guided Workflows: Systematic investigation phases prevent rushed analysis
  • Team Dynamics Under Control: CLAUDE might initiate analysis, then delegate subtasks to Gemini or O3, with each model having full visibility into prior discussions
  • Fresh Context Windows: Offload heavy tasks (code reviews, bug hunting) to isolated subagents while keeping your main session clean
  • True Provider Abstraction: Switch between or combine any AI provider without changing your workflow

Pricing

Free; source on GitHub (licence shown as “Other” by GitHub; earlier note said Apache-2.0, unconfirmed)

Requirements:

  • Python 3.10+
  • API credentials from your chosen providers (OpenAI, Anthropic, Google, etc.)
  • You pay for API usage from each provider separately

Getting Started

Quick Installation (5-minute setup):

git clone https://github.com/BeehiveInnovations/pal-mcp-server.git  
cd pal-mcp-server  
  
# Handles setup, config, API keys from environment  
# Auto-configures Claude Desktop, Claude Code, Gemini CLI, Codex CLI, Qwen CLI  
./run-server.sh  

The script handles:

  • Environment setup and dependency installation
  • API key configuration guidance
  • Auto-detection of common AI desktop clients
  • Configuration of Claude Desktop, Claude Code, and other tools

Configuration:

  • Configure via environment variables
  • Edit .env to enable/disable tools and preserve context window space
  • Supports sensible defaults for most use cases

Use Cases

  • Code Review: Spawn a fresh Gemini instance to review code while CLAUDE continues development
  • Multi-Perspective Debugging: Get different AI models to analyze the same bug from various angles
  • Security Audits: Coordinate multiple models to identify vulnerabilities
  • Collaborative Planning: Run debates between models to reach deeper insights
  • Large Codebase Analysis: Use Gemini’s 1M token context for massive projects
  • Documentation Generation: Coordinate models for comprehensive docs
  • Consensus Building: Get multiple AI opinions on architectural decisions

Technical Details

  • Protocol: Model Context Protocol (MCP) server
  • Language: Python 3.10+
  • Supported Providers: 50+ models via OpenAI, Anthropic, Google, Grok, Azure, Ollama, OpenRouter, custom endpoints
  • Context Management: Conversation threading with persistent state
  • Token Limits: Auto-handles large prompts to work within MCP constraints
  • Vision Models: Full support for image/diagram analysis

Integration

Works with:

Sources