- Documentation
- Integrations
- Apps
- Cohere integration
Cohere integration
Generate chat completions, embed text, rerank search results, parse documents, and transcribe audio with Cohere models on behalf of your clients.
What it does
The Cohere integration lets your agency call Cohere's Command, Embed, Rerank, Parse, and Transcribe models from inside any workflow. Connect a client's Cohere key once and your workflows can turn an emailed invoice image into structured markdown, transcribe a call recording into a CRM note, rerank retrieval candidates before handing them to another model, build a semantic index over a client's knowledge base, or extract structured JSON from unstructured text.
Rerank and Parse are the two capabilities clients most often cannot get elsewhere: Rerank scores candidate documents against a query far more accurately than embedding similarity alone, and Parse reads a document image into clean markdown or block structure without a separate OCR vendor.
Connect a Cohere account
- Open your workspace in TaskJuice and navigate to Connections.
- Choose Cohere and click Connect.
- In a new tab, open the Cohere API keys dashboard signed in as the client (or as your agency, if the client has delegated key creation to you).
- Click New Trial Key or New Production Key, give it a name that identifies the workspace, and copy the key value.
- Paste the key into TaskJuice and save the connection.
Cohere does not offer OAuth for API access, so a pasted key is the only supported method. Keys are unscoped: a key can reach every endpoint the client's organization is entitled to. To rotate or revoke one, delete the entry in the Cohere dashboard, generate a new key, and update the TaskJuice connection.
Triggers
Cohere publishes no webhook surface and no event API, so the integration is action-only. Use a Schedule trigger, another app's trigger, or an inbound webhook to start the workflow, then call Cohere inside it.
Actions
cohere/create-chat-completiongenerates a chat completion against a chosen Command model. Supports tool definitions, grounding documents with citations, structured JSON output via response format, reasoning with a token budget, safety mode, and the usual sampling controls.cohere/create-embeddingreturns one embedding per input using a chosen Embed model.embed-v4.0also accepts images and lets you pick an output dimension of 256, 512, 1024, or 1536.cohere/rerank-documentsreorders a list of documents by relevance to a query, optionally trimmed to the top N and truncated per document.cohere/parse-documentconverts a document image into markdown or structured blocks. Use it for invoices, receipts, contracts, and forms.cohere/transcribe-audiotranscribes an audio file to text. Accepts flac, mp3, mpeg, mpga, ogg, and wav.cohere/list-modelslists the models available to the connected account, optionally filtered to one endpoint.
Every action's Model field is a dropdown populated from the connected account, so you pick a model that key can actually reach rather than pasting an ID that may have been retired.
Known limitations
- Parse takes images, not PDFs. The endpoint accepts an http(s) image URL or a base64 data URI, capped at 20 MB and 50 megapixels. Convert PDF pages to images upstream before calling it.
- Transcribe is rate-limited to 5 requests per minute on trial keys, and production access for the transcription models is granted by Cohere's sales team rather than self-serve. Check the client's entitlement before building a high-volume transcription workflow.
- Trial keys are capped at 1,000 API calls per month across all endpoints, and are not licensed for production use. Move clients to a production key before launch.
- Cohere retires models on a published schedule. Models such as
command-r-plus,command-r, and the v2.0 embed and rerank families now return a 404 rather than a deprecation warning. The Model dropdown always reflects what the connected key can reach today; pin a dated model ID (for examplecommand-a-03-2025) rather than a floating alias when you need reproducibility. - Streaming responses are not exposed. Each chat completion runs as a single request and resolves only after the full response returns.
- Embedding requests accept up to 96 inputs per call. Split larger batches across a Loop node so each batch runs as its own request.
- Reasoning output arrives as a separate content block. When you enable thinking on a chat completion, read the assistant text by selecting the block whose
typeistext, not by taking the first element ofcontent. - Per-key rate limits, monthly quotas, and model entitlements are governed by the client's Cohere organization, not by TaskJuice. A 429 surfaces as a retryable rate-limit error with backoff.