Skip to main content

Cohere integration

Generate chat completions, embed text, rerank search results, parse documents, and transcribe audio with Cohere models on behalf of your clients.

What it does

The Cohere integration lets your agency call Cohere's Command, Embed, Rerank, Parse, and Transcribe models from inside any workflow. Connect a client's Cohere key once and your workflows can turn an emailed invoice image into structured markdown, transcribe a call recording into a CRM note, rerank retrieval candidates before handing them to another model, build a semantic index over a client's knowledge base, or extract structured JSON from unstructured text.

Rerank and Parse are the two capabilities clients most often cannot get elsewhere: Rerank scores candidate documents against a query far more accurately than embedding similarity alone, and Parse reads a document image into clean markdown or block structure without a separate OCR vendor.

Connect a Cohere account

  1. Open your workspace in TaskJuice and navigate to Connections.
  2. Choose Cohere and click Connect.
  3. In a new tab, open the Cohere API keys dashboard signed in as the client (or as your agency, if the client has delegated key creation to you).
  4. Click New Trial Key or New Production Key, give it a name that identifies the workspace, and copy the key value.
  5. Paste the key into TaskJuice and save the connection.

Cohere does not offer OAuth for API access, so a pasted key is the only supported method. Keys are unscoped: a key can reach every endpoint the client's organization is entitled to. To rotate or revoke one, delete the entry in the Cohere dashboard, generate a new key, and update the TaskJuice connection.

Triggers

Cohere publishes no webhook surface and no event API, so the integration is action-only. Use a Schedule trigger, another app's trigger, or an inbound webhook to start the workflow, then call Cohere inside it.

Actions

  • cohere/create-chat-completion generates a chat completion against a chosen Command model. Supports tool definitions, grounding documents with citations, structured JSON output via response format, reasoning with a token budget, safety mode, and the usual sampling controls.
  • cohere/create-embedding returns one embedding per input using a chosen Embed model. embed-v4.0 also accepts images and lets you pick an output dimension of 256, 512, 1024, or 1536.
  • cohere/rerank-documents reorders a list of documents by relevance to a query, optionally trimmed to the top N and truncated per document.
  • cohere/parse-document converts a document image into markdown or structured blocks. Use it for invoices, receipts, contracts, and forms.
  • cohere/transcribe-audio transcribes an audio file to text. Accepts flac, mp3, mpeg, mpga, ogg, and wav.
  • cohere/list-models lists the models available to the connected account, optionally filtered to one endpoint.

Every action's Model field is a dropdown populated from the connected account, so you pick a model that key can actually reach rather than pasting an ID that may have been retired.

Known limitations

  • Parse takes images, not PDFs. The endpoint accepts an http(s) image URL or a base64 data URI, capped at 20 MB and 50 megapixels. Convert PDF pages to images upstream before calling it.
  • Transcribe is rate-limited to 5 requests per minute on trial keys, and production access for the transcription models is granted by Cohere's sales team rather than self-serve. Check the client's entitlement before building a high-volume transcription workflow.
  • Trial keys are capped at 1,000 API calls per month across all endpoints, and are not licensed for production use. Move clients to a production key before launch.
  • Cohere retires models on a published schedule. Models such as command-r-plus, command-r, and the v2.0 embed and rerank families now return a 404 rather than a deprecation warning. The Model dropdown always reflects what the connected key can reach today; pin a dated model ID (for example command-a-03-2025) rather than a floating alias when you need reproducibility.
  • Streaming responses are not exposed. Each chat completion runs as a single request and resolves only after the full response returns.
  • Embedding requests accept up to 96 inputs per call. Split larger batches across a Loop node so each batch runs as its own request.
  • Reasoning output arrives as a separate content block. When you enable thinking on a chat completion, read the assistant text by selecting the block whose type is text, not by taking the first element of content.
  • Per-key rate limits, monthly quotas, and model entitlements are governed by the client's Cohere organization, not by TaskJuice. A 429 surfaces as a retryable rate-limit error with backoff.
Was this helpful?