Quick answer: Pick the integration surface first: it decides where your AI voice agent must stop talking. Genesys Audio Connector streams both ways but only in the IVR, with up to five integrations per org. Amazon Connect Customer’s A2A voice streaming carries PCM at 8, 16 or 24 kHz, with no mid-stream audio handback. To keep speaking after a human joins, use SIP transfer, the carrier or a Connect conference.
How do I put an AI voice agent in front of Genesys or Amazon Connect?
If you already run Genesys Cloud or Amazon Connect, the question is not “which AI voice agent” but “which door into the contact centre”. Each platform publishes several doors, and they differ in three ways that matter more than the model behind them: whether audio flows both ways, what format it arrives in, and the exact point at which the platform takes the caller back. The sequence below is the order that avoids rebuilding the integration after the first pilot.
- Day one: write down where the AI agent must stop talking. One sentence: “the agent hands to a human at X”, or “the agent stays on the call through the human conversation”. This single sentence picks your surface, and it is the step teams skip.
- Pick the integration surface from the four in the table below: bidirectional bot audio over the platform’s own websocket protocol, a listen-only media fork, a SIP transfer out and back, or an agent sitting in front of the contact centre at the carrier.
- Confirm the licence and deployment preconditions on the vendor’s own page. On Genesys, Audio Connector needs a Genesys Cloud CX 1, 2, 3 or 4 licence plus BYOT Rate E added to your subscription, and is not supported under BYOC Premises. On Amazon Connect, agent-to-agent (A2A) collaboration needs the Connect Customer tier and a Lex V2 bot.
- Build the agent endpoint to the platform’s audio contract: a websocket speaking the AudioHook Protocol for Genesys, an A2A websocket accepting PCM for Connect, or a SIP endpoint for the transfer surfaces.
- Wire the return path before the happy path. Decide what the agent sends back (output variables, an ESCALATE signal, a SIP transfer into a known entry point) and test the failure branch first.
- Finish state: a test call that enters the contact centre, reaches the AI agent, is handed to a human queue with the collected fields visible to the human, and a second test call where you kill the agent endpoint mid-sentence and the caller still reaches a person.
The first action on day one is that sentence in step 1, not an API key. Everything below is organised around it.
The four integration surfaces, and what each platform’s own docs permit
This is the threshold table the rest of the page relies on. Every cell comes from the vendor’s own documentation, read on 24 September 2026. Where a vendor page could not be read, the row says so instead of guessing.
| Surface | Genesys Cloud mechanism | Amazon Connect Customer mechanism | Audio direction | Documented format and limits |
|---|---|---|---|---|
| 1. Bidirectional bot audio over the platform websocket | Audio Connector, started by the Call Audio Connector action in Architect, over the AudioHook Protocol | A2A agent-to-agent collaboration with voice streaming (audioStreamingEnabled: true) |
Both ways | Genesys: up to 5 Audio Connector integrations; the action fails after 5 successful runs in one flow; Genesys’ reference server types media as PCMU or L16 at 8000 Hz. Connect: PCM at 8, 16 or 24 kHz, API-only setup, one active collaborator at a time |
| 1b. Turn-based bot (text in, text out) | No third-party voice equivalent: Genesys Bot Connector (now listed as legacy) takes up to 5 third-party bots and 60,000 bot invocations per minute per org, but only in Architect message flows. For turn-based voice, Genesys documents native integrations such as Amazon Lex V2 in Architect voice flows | A2A text streaming: Connect converts speech to text and speaks the reply with Connect Agentic voices | Text turns (on Connect, the platform owns speech) | Genesys: Bot Connector and its successor Digital Bot Connector are documented for message flows only. Connect: text mode is the default when the field is not set |
| 2. Listen-only media fork | AudioHook Monitor | Live media streaming to Kinesis Video Streams (Start media streaming block) | From the call to you | Genesys: up to 20 integrations, a maximum of 5 streaming at once, not for PCI secure flows. Connect: 8 kHz, one Kinesis video stream per active call, caller and heard audio on separate tracks |
| 3. SIP transfer out and back | BYOC Cloud SIP trunk from the Genesys Media Tier in AWS to your endpoint | Transfer to phone number block (PSTN), or the gated CONNECT_CALL_TRANSFER_CONNECTOR through an Amazon Chime SDK Voice Connector |
Both ways; the AI agent owns the media until it transfers back | Genesys: RTP over the public internet, your endpoint needs a public IP. Connect PSTN: “Resume flow after disconnect” works only if the external party hangs up; calls to countries outside Connect’s default list fail until AWS approves a quota increase request. Connect SIP connector: gated, via your AWS account team |
| 4. Agent in front, at the carrier | Neither platform is in the path until your AI agent transfers the caller in. Your own SBC or trunk terminates the call on the agent first. | Both ways, entirely yours | Bounded by your trunk and SBC, not by the contact-centre platform | |
The one-line reading of that table: Genesys and Amazon Connect both now publish a native, bidirectional way to put an external AI voice agent on a live call, and both end that agent’s turn at a platform-controlled point.
The Containment Boundary: where the AI agent’s turn ends on each surface
The Containment Boundary is the point in a call after which your AI voice agent is no longer allowed in the media path. It is a property of the integration surface, not of the agent, and it is the thing buyers most often discover after the pilot. The decision rule: if the AI agent must still be speaking after your contact centre hands the caller to a human, a platform-native bot socket is the wrong surface; use a SIP transfer out and back, put the agent in front at the carrier, or, on Connect, have the human agent conference it in as an external party. If it only needs to listen past that point, add a listen-only fork.
| Surface | Where the Containment Boundary sits | What the vendor’s own page says | Use it when |
|---|---|---|---|
| Genesys Audio Connector | End of the IVR leg | “The bi-directional streaming session is active only in the IVR channel. It does not transfer to an agent.” | The agent finishes before a human picks up: qualification, authentication, routing, booking |
| Connect A2A voice streaming | When the collaborator sends COMPLETE or ESCALATE (AWS lists audio handback under “Features not yet available”, so this may change) | “Audio handback is not supported. In this release, once a collaborator takes over a voice contact in voice streaming mode, audio control cannot be returned mid-stream to your Connect AI agent while the A2A session remains active.” | The external agent owns a self-contained task, then control returns to the Connect AI agent |
| Genesys AudioHook Monitor | No speaking role; listening can continue into the agent leg | Lists monitoring “customer participants during the AudioHook sessions as they converse with other participants (for example, with a bot or agent)” | The AI needs to hear the human conversation (assist, notes), not speak in it |
| Connect live media streaming | No speaking role; capture continues until a Stop media streaming block | “Customer audio is captured until a Stop media streaming block is invoked, even if the contact is passed to another flow.” | Same as above, on Connect |
| SIP transfer out and back | Wherever your agent transfers the caller back | Connect: “Resume flow after disconnect” returns the caller “and resumes the flow after the transferred call ends” | The agent must run a long conversation the IVR-only socket cannot hold |
| Agent at the carrier | Wherever you put it | Not a platform feature; your design | The AI answers first and the contact centre is the escalation target |
A worked reading: an outbound qualification agent that books a meeting and ends the call sits comfortably inside Audio Connector’s boundary. An inbound agent that must stay on as a live interpreter while a human talks does not fit Audio Connector or A2A voice streaming at all, because on both platforms the speaking session ends before or at the human. On Amazon Connect the documented route for that case is a conference: a human agent can use quick connects or the number pad to add an external party to a live call, and AWS’s own example of such a participant is a translator, so an AI interpreter reachable on a phone number can join after the human does.
Genesys Cloud: setting up Audio Connector for an external voice bot
Genesys describes Audio Connector as “a mechanism and generic protocol to provide a bi-directional and near real-time stream of voice interactions from the Genesys Cloud platform to a third-party voice bot provider and back”. The setup order Genesys documents is: obtain the Audio Connector integration from Genesys AppFoundry, configure and activate it, then call it from Architect with the Call Audio Connector action.
- Commercials first. The Audio Connector pricing page says you “must contact Genesys Cloud Sales to add BYOT Rate E to your subscription” before you can obtain the premium application, and that Genesys “charges for the Audio Connector session when streaming begins and stops charging when streaming ends”, rounded to the nearest minute at month end.
- Check your edge. Audio Connector “is not supported under BYOC Premises” and “is not supported for premises-based Edge (LDM)”. If your Genesys telephony is on premises, surfaces 1 and 2 are closed to you (AudioHook Monitor is not supported there either) and surface 3 or 4 is the route.
- Stand up the server. Genesys links a reference server, the AudioConnector server reference implementation on GitHub. Its protocol types define the media format as PCMU or L16 and the rate as 8000, so plan for 8 kHz telephony audio into your speech recogniser. Genesys also advises creating Audio Connector servers “in the same or near region” because the stream crosses the internet.
- Place the action. The Call Audio Connector action is available “in inbound, outbound, secure, and in-queue call flows”. Genesys appends your Connector ID to the integration’s Base Connection URI when it opens the websocket, so one server can host several bots.
- Map inputs and outputs. Pass caller context in as input variables. When the bot completes, “Architect assigns the key/value pairs that your Audio Connector integration returns to the output variables” you defined. This is how what the agent learned reaches the human: as flow variables, not as a live voice.
- Design around the five-run ceiling. The action’s failure outputs include
ActionInvocationLimitExceeded: “If a flow has successfully run the Call Audio Connector action five times, any subsequent invocations of the action take the failure path.” An IVR that bounces the caller back to the bot after every sub-task hits this on the sixth bounce.
Two further limits shape capacity. Genesys caps an organisation at five Audio Connector integrations, with more requested through the Genesys Cloud Ideas Portal, and “Audio Connector supports one bi-directional stream” per session. The quotable version: Genesys Audio Connector gives an external AI voice agent a full-duplex seat in the IVR, and only in the IVR.
Amazon Connect Customer: A2A voice streaming for an external AI agent
Two facts changed recently, and both matter if you are following an older guide. First, the product has been renamed: AWS’s administrator guide now opens with “Amazon Connect Customer is the current name for the product previously called Amazon Connect.” Second, the native route for an external AI voice agent is now agent-to-agent collaboration over the A2A protocol, which AWS describes as “an open standard originally developed by Google and now governed by the Linux Foundation”, with Connect-specific extensions for voice streaming.
- Check the tier. “Agent-to-agent collaboration is not available on Customer Basic.” Your instance must be on the Connect Customer tier.
- Add the Lex V2 bot. Voice contacts require “a Lex V2 bot configured with the AMAZON.QInConnectIntent”. The Connect AI agent is the orchestrator; your external agent is a collaborator it brings in.
- Expose an A2A websocket. The external agent “must support the A2A protocol over a WebSocket endpoint”, authenticated with an API key that Connect presents as a bearer token on the WebSocket upgrade.
- Configure by API, not console. This is “an API-only feature” and “all setup requests must be SigV4-signed”. AWS notes that the installed CLI “does not yet include the newest shapes”, so some steps use awscurl.
- Choose the voice mode per collaborator. Set
audioStreamingEnabledonhandoffAgentConfiguration:false(the default) for text streaming, where Connect handles speech, ortruefor voice streaming, where your agent runs its own recognition and voice. Voice streaming “uses PCM audio encoding” at “8 kHz, 16 kHz, and 24 kHz”, negotiated at session start. - Send trace data. External collaborators “are required to send trace data back” to Connect, and an endpoint that does not “will be disabled for new contacts”. Build this before go-live, not after.
- Validate latency with AWS. For voice streaming, AWS asks you to “work with your AWS account team or Solutions Architect to validate that the agent meets the latency and audio quality requirements before deploying to production”.
The collaboration ends when your agent signals Completed or Escalate, or through a fallback path if it stops responding. Escalate is the signal that “the contact needs a human agent”. The quotable version: on Amazon Connect Customer, an external AI voice agent speaks as a guest of the Connect AI agent and leaves the call when it signals Completed or Escalate.
What happened to Amazon Connect external voice transfer?
Older guides, including community answers in an AWS re:Post thread on integrating an external voicebot with Amazon Connect without PSTN, point to AWS’s documentation on the external voice transfer connector. Its admin-guide page, “Set up Amazon Connect external voice transfer to an on-premise voice system”, described moving calls between Amazon Connect and other voice systems “without using the PSTN”. On 24 September 2026 that URL redirects to the guide’s landing page, the page is absent from the current guide’s table of contents, and the service quotas page no longer lists an external voice transfer connector quota. The last archived copy we could read is dated 8 December 2025.
The capability has not simply vanished. The Amazon Chime SDK API reference still lists CONNECT_CALL_TRANSFER_CONNECTOR as an integration type that lets enterprises “directly transfer voice calls and metadata without using the public telephone network”, and adds: “This integration is a gated feature. Please reach out to your account team to discuss this feature with a Connect Specialist.” It lists a second type, CONNECT_ANALYTICS_CONNECTOR, for sending audio from other voice systems into Connect’s analytics, which is the reverse direction and not a way to put an agent in front. Treat the SIP connector as an account-team conversation, not a self-serve setting.
Listen-only forks: when the AI agent only needs to hear the call
Not every AI voice agent needs a mouth. Call summaries, compliance flags and live agent assist only need audio in, and both platforms publish a fork for that which, unlike the bot sockets, can keep listening after a human picks up.
- Genesys AudioHook Monitor streams “voice interactions from the Genesys Cloud platform to any third-party service endpoint”. Genesys supports “up to 20 AudioHook Monitor integrations” but “can only stream audio to a maximum of five active integrations at a given time”, and Genesys says that limit cannot be changed. Genesys recommends against using it during PCI-relevant interactions such as Architect secure flows.
- Amazon Connect live media streaming sends “all audio to and from the customer to Kinesis Video Streams”, with what the customer says “on a separate track from what the customer hears”, at “a sampling rate of 8 kHz”. Connect uses “one Kinesis video stream” per active call, so a large deployment may need to ask AWS Support for a higher Kinesis Video Streams quota.
A common pattern is to pair the two surfaces: a bot socket for the IVR leg, and a monitor fork that follows the caller into the human conversation so the AI’s notes keep updating.
SIP transfer out and back: what breaks first
The SIP route gives your AI agent the whole conversation, and moves three problems onto you. Our guide to connecting an AI agent to your own SIP trunk covers trunk configuration in depth; these are the problems specific to sitting behind a contact centre.
The trunk identity is not the dialled number. In a public LiveKit issue from October 2025, an operator reported that inbound INVITEs from a Genesys trunk were rejected with a 486 and a “flood” reason, while the same setup worked with Twilio. The workaround the reporter later posted in the thread, which was closed as not planned: with Genesys-defined trunks, “the number you need to use is actually not the CPN but rather the internal routed id”. If your SIP server matches calls on the called number, expect to map the Genesys routing identifier instead.
The public internet is in the media path. Genesys BYOC Cloud “involves Real-time Transport Protocol traffic traveling over the Internet” between Genesys’s AWS resources and your trunk endpoint, which “must have a publicly reachable IP address”. If you would rather run the agent on your own SBC with no CPaaS in between, that is the pattern in our guide to running an AI voice agent on direct SIP from your own SBC; this page starts where that one ends, inside someone else’s contact centre.
The return trip loses context unless you carry it. On Amazon Connect, the Transfer to phone number block can resume the flow after the external call ends, but that “works only if the external party disconnects, and the customer doesn’t disconnect”. That is a usable pattern: the AI agent hangs up, the caller drops back into the Connect flow, and the flow queues them for a human. What the agent learned has to come back through your own API or contact attributes, because the audio leg carries none of it. Our write-up of AI-to-human handoff and context transfer covers the field-by-field version, and the AI voice agent transfer failure checklist covers the ways the transfer itself fails.
Which surface should I choose? The threshold table
| If this is true | Use | Do not use |
|---|---|---|
| Agent finishes before any human, and you are on Genesys Cloud (cloud edge) | Audio Connector | SIP transfer, which adds a trunk you do not need |
| Agent finishes before any human, Genesys on BYOC Premises or premises Edge | SIP transfer out and back, or agent at the carrier | Audio Connector or AudioHook Monitor (both unsupported there) |
| Agent needs more than 5 bot legs in one Genesys flow | Fewer, longer bot sessions, or SIP transfer | Looping back to Call Audio Connector (the sixth invocation takes the failure path) |
| On Amazon Connect Customer tier, agent owns a self-contained task | A2A voice streaming | PSTN transfer, which AWS’s archived comparison says cannot carry call metadata (the current block documents only Send DTMF and caller ID settings) |
| On Customer Basic | Transfer to phone number, or agent at the carrier | A2A collaboration (not available on Customer Basic) |
| Agent must keep speaking after a human joins | Agent at the carrier, SIP transfer where the agent keeps the call, or (on Connect) the human agent adding the AI agent’s number to the call as an external party | Audio Connector, A2A voice streaming |
| Agent only needs to hear the human conversation | AudioHook Monitor or live media streaming | A bidirectional bot socket |
What running it yourself costs, beyond the platform fee
Every surface above moves audio. None of them decides what the AI voice agent says. On the platform side you pay per streaming minute (Genesys BYOT Rate E) or per collaboration mode (Connect prices A2A by text, unidirectional voice and bidirectional voice). On your side you run a websocket or SIP service that must stay up whenever the contact centre is open, a speech stack tuned for 8 kHz telephony audio, trace export on Connect, a return path tested for failure, and someone who owns the conversation logic week after week. The integration gets finished; the conversation logic never is.
That last part is where the outcome is decided. Zian has run outbound acquisition since 2017. Its PrecisionPitch AI™ continuously split-tests scripts and approaches, its learning engine tracks around 420,000 data points across more than 10,000 leads a day, and it has taken accounts from roughly 2% conversion to around 8%. Zian does not publish a Genesys or Amazon Connect connector (its published integrations are on the Zian features and integrations page), so the surface choice above is yours to make with your contact-centre team either way.
Frequently asked questions
Can I integrate an external voicebot with Amazon Connect without the PSTN?
Yes. AWS currently documents two routes. The one you configure yourself through the API is agent-to-agent collaboration with voice streaming, where Connect streams the caller’s audio to your agent over an A2A websocket. The SIP route is the CONNECT_CALL_TRANSFER_CONNECTOR integration type, which the Amazon Chime SDK CreateVoiceConnector API reference describes as a gated feature you request through your AWS account team.
Can a Genesys Audio Connector bot stay on the call after the transfer to an agent?
No. Genesys states that the bi-directional streaming session is active only in the IVR channel and does not transfer to an agent. The bot can return key and value pairs to Architect as output variables, and standard Architect transfer behaviour applies after the session ends.
How many Audio Connector integrations can a Genesys Cloud org have?
Five, per the Genesys Audio Connector overview. More can be requested through the Genesys Cloud Ideas Portal. Separately, a single flow that successfully runs the Call Audio Connector action five times sends any further invocation down the failure path.
What audio format does Amazon Connect send to an external AI agent?
In A2A voice streaming mode, PCM audio at 8 kHz, 16 kHz or 24 kHz, negotiated when the session starts, per the Connect voice configuration page for collaborating AI agents. The listen-only Kinesis Video Streams route uses 8 kHz.
Is Amazon Connect now called Amazon Connect Customer?
Yes. The AWS administrator guide says Amazon Connect Customer is the current name for the product previously called Amazon Connect, and that Amazon Connect is now a set of agentic AI solutions for different business functions.
Why does my SIP server reject INVITEs from a Genesys trunk as a flood?
In one public LiveKit SIP issue, the workaround the reporter posted was to configure the trunk with the Genesys internal routed identifier instead of the phone number (the thread calls it the CPN). Check which number your SIP server matches on before you look at rate limits.
Does Genesys Audio Connector work with BYOC Premises?
No. Genesys lists Audio Connector as not supported under BYOC Premises or premises-based Edge. On those deployments, use a SIP transfer or put the agent in front of Genesys at the carrier.
Where every figure on this page comes from
| Figure | Who published it | Link | Date read |
|---|---|---|---|
| Audio Connector: IVR channel only, does not transfer to an agent; up to 5 integrations; one bi-directional stream; not supported under BYOC Premises or premises Edge | Genesys | help.genesys.cloud/articles/audio-connector-overview/ | 2026-09-24 |
| Call Audio Connector action: inbound, outbound, secure and in-queue call flows; failure path after 5 successful runs | Genesys | help.genesys.cloud/articles/call-audio-connector-action/ | 2026-09-24 |
| Genesys Cloud CX 1, 2, 3 or 4 licence and BYOT Rate E required; billed while streaming, rounded to the nearest minute at month end | Genesys | help.genesys.cloud/articles/audio-connector-pricing/ | 2026-09-24 |
| Media format PCMU or L16, rate 8000 | Genesys (GenesysCloudBlueprints reference server) | github.com/GenesysCloudBlueprints/audioconnector-server-reference-implementation (src/protocol/core.ts) | 2026-09-24 |
| AudioHook Monitor: up to 20 integrations, 5 streaming at once; not for PCI secure flows; not supported under BYOC Premises or premises Edge | Genesys | help.genesys.cloud/articles/audiohook-monitor-overview/ | 2026-09-24 |
| Bot Connector: up to 5 third-party bot integrations; Architect message flows | Genesys | help.genesys.cloud/articles/about-genesys-bot-connector/ | 2026-09-24 |
| Bot Connector listed as legacy; Digital Bot Connector for message flows; Amazon Lex V2 in Architect voice flows | Genesys | help.genesys.cloud/articles/about-bots/ | 2026-09-24 |
| Bot Connector: 60,000 bot invocations per minute per org | Genesys | help.genesys.cloud/faqs/are-there-limits-on-the-number-of-concurrent-bot-connector-interactions/ | 2026-09-24 |
| BYOC Cloud: RTP over the internet; public IP required | Genesys | help.genesys.cloud/articles/byoc-cloud-overview/ | 2026-09-24 |
| Amazon Connect renamed Amazon Connect Customer | AWS | docs.aws.amazon.com/connect/latest/adminguide/what-is-amazon-connect.html | 2026-09-24 |
| A2A: PCM at 8, 16 and 24 kHz; audio handback not supported; text mode default | AWS | docs.aws.amazon.com/connect/latest/adminguide/a2a-voice.html | 2026-09-24 |
| A2A: Connect Customer tier, Lex V2 bot with AMAZON.QInConnectIntent, API-only, one active collaborator, priced by collaboration mode; audio handback listed under features not yet available | AWS | docs.aws.amazon.com/connect/latest/adminguide/a2a-quotas.html | 2026-09-24 |
| A2A: websocket, bearer token, SigV4, awscurl, latency validation with AWS | AWS | docs.aws.amazon.com/connect/latest/adminguide/a2a-setup-external.html | 2026-09-24 |
| A2A: Completed and Escalate outcomes; fallback path | AWS | docs.aws.amazon.com/connect/latest/adminguide/a2a-how-it-works.html | 2026-09-24 |
| A2A: external collaborators “required to send trace data back”; endpoint “disabled for new contacts” without it | AWS | docs.aws.amazon.com/connect/latest/adminguide/a2a-collaboration.html | 2026-09-24 |
| Agents can add an external party to a live call via quick connects or the number pad; translator given as an example participant | AWS | docs.aws.amazon.com/connect/latest/adminguide/multi-party-calls.html | 2026-09-24 |
| Live media streaming: 8 kHz; one Kinesis video stream per active call; separate tracks | AWS | docs.aws.amazon.com/connect/latest/adminguide/plan-live-media-streams.html | 2026-09-24 |
| Capture continues until a Stop media streaming block | AWS | docs.aws.amazon.com/connect/latest/adminguide/start-media-streaming.html | 2026-09-24 |
| Resume flow after disconnect works only if the external party disconnects; outbound country allowlist; Send DTMF and caller ID settings | AWS | docs.aws.amazon.com/connect/latest/adminguide/transfer-to-phone-number.html | 2026-09-24 |
| CONNECT_CALL_TRANSFER_CONNECTOR is a gated feature; CONNECT_ANALYTICS_CONNECTOR | AWS | docs.aws.amazon.com/chime-sdk/latest/APIReference/API_voice-chime_CreateVoiceConnector.html | 2026-09-24 |
| External voice transfer page, last archived copy 8 December 2025; its comparison table says metadata “Cannot be transferred” over a transfer to phone number | AWS (archived by the Internet Archive) | web.archive.org snapshot of external-voice-transfer.html | 2026-09-24 |
| Genesys trunk INVITEs rejected with 486 “flood”; reporter’s workaround was the internal routed id; closed as not planned | LiveKit SIP issue tracker (operator report, October 2025) | github.com/livekit/sip/issues/501 | 2026-09-24 |
| Zian: since 2017; around 420,000 data points; more than 10,000 leads a day; roughly 2% to around 8% | Zian AI (first-party) | zian.ai | 2026-09-24 |