diff --git a/README.md b/README.md index ae363c4..c6e3df9 100644 --- a/README.md +++ b/README.md @@ -18,6 +18,8 @@ from MCP servers (Yahoo Finance, Sharesight) and web search. It runs on **xAI (G 1. Create a bot with @BotFather. Run `/setprivacy` and choose **Disable**, so the bot sees every message in a group (remove and re-add it to groups it is already in). Who may use the bot is controlled in Telegram/BotFather; the bot itself has no allow-list. + - **IMPORTANT:** if the bot has access to Sharesight, enable "Restrict bot usage" in BotFather (this + enables user whitelisting). 2. Python 3.13.5 (production; GitHub CI runs the latest stable Python and Node.js as an early warning): `pip install -r requirements.txt` (plus Node.js for npx-based MCP servers such as Yahoo Finance). 3. Export `TELEGRAM_BOT_TOKEN` and **exactly one** of `XAI_API_KEY`, `OPENROUTER_API_KEY` or `ZAI_API_KEY`. 4. Optional: edit `mcp_servers.json` (see [MCP data tools](#mcp-data-tools)). @@ -239,18 +241,28 @@ token size of the tool definitions. The `Starting ...` line shows the git commit Figures from search results and Artificial Analysis at the time of writing (October 2026), so treat them as rough; speed varies by provider. -| Model | $ per 1M in / out | Output speed | Quality index | -|---|---|---|---| -| `xiaomi/mimo-v2.6-pro` | 0.43 / 0.87 | ~28-46 tok/s | 46 (top open-weight model) | -| `xiaomi/mimo-v2.6-flash` | 0.14 / 0.28 | ~56 tok/s | 38 | -| `z-ai/glm-5.3-flash` (default on OpenRouter) | 0.15 / 0.50 | ~50 tok/s (other hosts up to ~270) | 57 | -| `grok-4.3` / `x-ai/grok-4.3` (default on xAI) | 1.25 / 2.50 | ~105-146 tok/s | 25 (at high reasoning) | -| `glm-5.3-flash` (default on z.ai) | 0.15 / 0.50 (cached input 0.03) | ~50 tok/s | 57 | -| `xiaomi/mimo-v2.5-pro` | 0.30 / 0.61 | ~29-46 tok/s | unreliable (retires 21 Oct 2026) | -| `xiaomi/mimo-v2.5` | 0.12 / 0.24 | ~44-58 tok/s | not found (retires 21 Oct 2026) | +| Model | $ per 1M in / out | Output speed | Quality index | Fast sibling for `FAST_MODEL` | +|---|---|---|---|---| +| `xiaomi/mimo-v2.6-pro` | 0.43 / 0.87 | ~28-46 tok/s | 46 (top open-weight model) | `xiaomi/mimo-v2.6-flash` | +| `xiaomi/mimo-v2.6-flash` | 0.14 / 0.28 | ~56 tok/s | 38 | itself (a flash model) | +| `z-ai/glm-5.3-flash` (default on OpenRouter) | 0.15 / 0.50 | ~50 tok/s (other hosts up to ~270) | 57 | itself (a flash model) | +| `grok-4.3` / `x-ai/grok-4.3` (default on xAI) | 1.25 / 2.50 | ~105-146 tok/s | 25 (at high reasoning) | none (see below) | +| `glm-5.3-flash` (default on z.ai) | 0.15 / 0.50 (cached input 0.03) | ~50 tok/s | 57 | itself (a flash model) | +| `glm-5.3-flashx` (z.ai) | 0.37 / 1.25 (cached input 0.075) | not found | not found | itself (a flash model) | +| `glm-5.3` (z.ai) | 1.40 / 4.40 (cached input 0.26) | not found | not found | `glm-5.3-flash` | +| `glm-5.2` (z.ai) | 1.40 / 4.40 (cached input 0.26) | not found | not found | `glm-5.3-flash` | +| `xiaomi/mimo-v2.5-pro` | 0.30 / 0.61 | ~29-46 tok/s | unreliable (retires 21 Oct 2026) | `xiaomi/mimo-v2.5` (also retiring) | +| `xiaomi/mimo-v2.5` | 0.12 / 0.24 | ~44-58 tok/s | not found (retires 21 Oct 2026) | itself | Most of the delay is reasoning time before the first answer token, not typing speed: use -`REASONING=low`, and `FAST_MODEL=xiaomi/mimo-v2.6-flash` so simple messages skip the big model. +`REASONING=low`, and set `FAST_MODEL` so simple messages skip the big model. `FAST_MODEL` runs on the +same provider as your API key: +- **OpenRouter:** any model, for example `FAST_MODEL=xiaomi/mimo-v2.6-flash` or `z-ai/glm-5.3-flash` + alongside a bigger `MODEL`. +- **z.ai:** a GLM model: for example `MODEL=glm-5.3` with `FAST_MODEL=glm-5.3-flash` (or `glm-5.3-flashx`). + A z.ai model missing from the price table in `lib/llm/zai.py` shows $0 cost. +- **xAI:** a Grok model, and there is no fast one to pick: Grok 4 Fast and Grok 4.1 Fast were reportedly + retired on 15 May 2026 and now redirect to `grok-4.3`. ## Development ``` diff --git a/lib/llm/zai.py b/lib/llm/zai.py index a7d2144..8d64005 100644 --- a/lib/llm/zai.py +++ b/lib/llm/zai.py @@ -16,7 +16,12 @@ BASE_URL = "https://api.z.ai/api/paas/v4/" # USD per million tokens (input, cached input, output); a model missing here shows $0. -PRICES = {"glm-5.3-flash": (0.15, 0.03, 0.50)} +PRICES = { + "glm-5.3-flash": (0.15, 0.03, 0.50), + "glm-5.3-flashx": (0.37, 0.075, 1.25), + "glm-5.3": (1.40, 0.26, 4.40), + "glm-5.2": (1.40, 0.26, 4.40), +} # z.ai models always think and accept only low, high or max. Low costs almost no reasoning # tokens, so it is the default; the bot's REASONING values map onto z.ai's. EFFORT = {"": "low", "low": "low", "medium": "high", "high": "high", "max": "max"} diff --git a/tests/test_llm_backends.py b/tests/test_llm_backends.py index c364cfc..91a7623 100644 --- a/tests/test_llm_backends.py +++ b/tests/test_llm_backends.py @@ -136,6 +136,18 @@ async def test_zai_reasoning_maps_to_its_efforts_and_none_leaves_tools_out(): assert client.kwargs[1]["extra_body"] == {"reasoning_effort": "max"} and "tools" not in client.kwargs[1] +def test_zai_prices_cover_the_published_glm_models(): + from types import SimpleNamespace as N + + from lib.llm.zai import PRICES + assert set(PRICES) == {"glm-5.3-flash", "glm-5.3-flashx", "glm-5.3", "glm-5.2"} + b = ZaiBackend(config.load(ZAI_ENV), FakeChatClient()) + usage = N(prompt_tokens=1_000_000, completion_tokens=1_000_000, prompt_tokens_details=N(cached_tokens=500_000)) + # 500k uncached in + 500k cached in + 1M out, per million tokens + assert b._cost(usage, "glm-5.3") == pytest.approx(0.5 * 1.40 + 0.5 * 0.26 + 4.40) + assert b._cost(usage, "glm-5.3-flashx") == pytest.approx(0.5 * 0.37 + 0.5 * 0.075 + 1.25) + + async def test_zai_warns_once_about_models_without_a_price(caplog): ZaiBackend(config.load({**ZAI_ENV, "MODEL": "glm-x", "FAST_MODEL": "glm-5.3-flash"}), FakeChatClient()) assert [r.getMessage() for r in caplog.records if "No z.ai price" in r.getMessage()] == [