JigsawStack Speech To Text
Convert Speech to Text using JigsawStack's AI Models.
Convert Speech to Text using JigsawStack's AI Models.
Text to speech for MCP clients. 23 languages, six voices, every render watermarked.
Kokoro Text to Speech (TTS) MCP Server
Text-to-speech API for AI agents. Convert text to natural speech audio in 20+ languages via Google TTS engine. Returns base64-encoded MP3 ready for playback. Tools: media_text_to_speech. Use this for voice notifications, accessibility features, audio content generation, or building voice interfaces. Returns: {audio (base64 MP3), language, duration}. No API key required — x402 micropayment $0.005/call on Base L2.
Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.
Kurdish (Sorani & Kurmanji) text-to-speech & speech-to-text — 664 AI voices. API key required.
Generate highly realistic Text to Speech voiceovers.
Convert Speech to Text using JigsawStack AI Models.
Convert text to MP3 speech audio in 50+ languages. x402 micropayment.
Bambara AI over MCP: text-to-speech, transcription and translation (Bamanankan + more).
Quote-first, non-custodial x402 text-to-speech with spend policy, MP3 artifacts, and receipts.
Frenchie is an MCP-first multimodal utility for AI agents. It helps agents work with files through a hosted MCP server: - OCR PDFs and images into clean Markdown - Transcribe audio and video into Markdown - Generate images from text prompts - More multimodal capabilities are coming soon Frenchie is built for agent workflows, not as a general OCR or speech-to-text platform. Files are processed for the requested job and results expire after delivery. Use it when your agent n
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
Voice Audio MCP server. Tools: text to speech, list voices, transcribe. Built by MEOK AI Labs.
Create and edit high-quality media across various formats including images, videos, audio, and 3D models. Enhance creative workflows by performing advanced tasks like background removal, image upscaling, and text-to-speech generation. Access a wide range of state-of-the-art generative models to bring complex visual and auditory ideas to life instantly.
Pay-per-call generative AI — 60+ models (image, video, music, speech). No signup, no API key; pay per call from your wallet over MPP/x402.
Stop recording your voice. NarrateAI adds professional voiceover to silent screen recordings — automatically. Just drop a video URL and get a narrated demo, dubbed video, or polished document back. Works from Cursor, Claude, ChatGPT, or any MCP client. **What it does:** - Narrates silent videos with AI voiceover — no mic, no script writing, no editing - 7 AI voices + voice cloning from a 15-second audio sample - Transcribes speech-to-text from any video (meetings, podcasts,
The API runtime for AI agents. One tool call. Any API. No setup. Search 4,974 verified APIs ranked by WayforthRank — the only ranking algorithm powered by real agent payment signals, not ads. Execute via 18 managed services with zero API keys. Self-healing: if a service fails, credits restore automatically. What your agent can do: Search 4,974 APIs across 19 categories by intent Execute: inference, translation, image generation, speech-to-text, web search, financial data, we
# Podstow **Your personal listen-later podcast.** Podstow turns the things you mean to read into a private podcast you can actually listen to. Connect it to Claude, ChatGPT, or any MCP-aware assistant and it can send web articles — or text the assistant writes itself (summaries, deep-research reports, digests) — straight into your personal podcast feed. Podstow scrapes the page (or ingests the text), runs it through natural text-to-speech, and publishes an episode to a priv
<a href="https://modelrunner.ai">ModelRunner</a> is a hosted remote MCP (Model Context Protocol) server that lets AI assistants like Claude and Cursor run 100+ AI models — text-to-image, image-to-image, text-to-video, image-to-video, video-to-video, music generation, speech-to-text, and image-to-3D, including Kling, HiDream, Hunyuan Image, Stable Audio, and Rodin. One connection exposes every model as a callable tool: search the catalog, inspect a model's input schema, run in
MCP server for local speech translation (EN ↔ 中文) via Whisper + Claude + Piper
Deterministic pay-per-call tools for AI agents: live web search and cited answers, news, browser render, market data, speech-to-text, wallet-keyed memory, plus 500+ long-tail tools via find_tool. Settle in USDC via x402 or free via proof-of-work. Maintained by Havok Holdings LLC.
One API for 100+ AI video, image, music and speech models.
Convert topics or existing files into professional presentation videos with automated slides and narration. Streamline content creation by generating outlines, scripts, and high-quality text-to-speech audio in a single workflow. Manage the entire production pipeline from initial rendering to final video export for polished results.
Analyze spoken language to provide immediate feedback on pronunciation accuracy and fluency. Identify specific phonetic errors to help learners improve their speaking skills in real-time. Guide users toward native-like speech through detailed scoring and actionable insights.
One key, 100+ models — chat with any LLM and generate video, images, speech. Free trial at 370.ai.
Reads text aloud locally on Windows, macOS, and Linux. No API key, account, or cloud service.
Pay-per-call AI: 59 generative models (image, video, music, speech). No signup, no API key.
Enable AI video generation, replica management, conversational AI, lipsync, and speech synthesis through a comprehensive MCP server for the Tavus API. Manage AI replicas, generate videos, create interactive conversations, synchronize audio, and generate speech seamlessly within MCP-compatible applications. Access full Tavus API v2 capabilities with robust error handling and easy integration.
Search speech in podcasts, government meetings, and your own audio: speakers, entities, timestamps.
Turn any video into publish-ready content using AI agents. 10 tools: create viral clips, generate animated captions in 100+ languages, dub videos in 80+ languages, translate subtitles, transcribe speech, reframe for Shorts/Reels/TikTok, apply brand templates, export videos, track processing, and retrieve results. Automate video repurposing, localization, and publishing workflows through natural conversation. Works with Claude Desktop, Claude Code, Cursor, VS Code, Windsurf,
Remote MCP for ViewMax AI video, image, music, and speech generation. Connect with Claude OAuth or a ViewMax API key (Authorization Bearer).
Create, inspect, and manage Wubble music, speech, voice, and sound-effect requests through MCP.
Universal MCP server giving any client (Claude Desktop, claude.ai, Cursor, Cline, VS Code) native access to the full Replicate catalog: image, video, music, speech (TTS+STT), LLMs, vision, upscale, inpaint, segment, embeddings, voice cloning, 3D, and lipsync. 29 tools, 63 curated models. Adds async batch jobs, DAG pipelines, a model recommender, a pre-call cost estimator, prediction history/cancel, multi-token round-robin, and both stdio and HTTP/SSE transports.
Generate AI images, video, speech and music inside Claude, ChatGPT, Cursor and any MCP client. 60+ models via one remote Streamable HTTP endpoint — credit-based pricing, sandbox keys for CI, no SDK required.
Manage Speko voice-AI agents, sessions, calls, phone numbers, knowledge bases, evals, and docs.