Vocalize: Lead Acquisition and User Engagement through Gamified Voice Competitions
Vocalize is a gamified voice-based WhatsApp competition system for lead acquisition and user engagement, evaluated at four live events in 2024.

Vocalize: Lead Acquisition and User Engagement through Gamified Voice Competitions

Authors

    Edvin Teskeredzic — Infobip, Sarajevo, Bosnia — Edvin.Teskeredzic@infobip.com

    Muamer Paric — Infobip, Sarajevo, Bosnia — Muamer.Paric@infobip.com

    Adna Sestic — Infobip, Sarajevo, Bosnia — Adna.Sestic@infobip.com

    Petra Fribert — Infobip, Zagreb, Croatia — Petra.Fribert@infobip.com

    Anamarija Lukac — Infobip, Zagreb, Croatia — Anamarija.Lukac@infobip.com

    Hadzem Hadzic — Infobip, Sarajevo, Bosnia — hadzem.hadzic@infobip.com

    Kemal Altwlkany — Infobip, Sarajevo, Bosnia — kemal.altwlkany@infobip.com

    Emanuel Lacic — Infobip, Zagreb, Croatia — emanuel.lacic@infobip.com

Publication

Proceedings of the 2025 Adjunct Proceedings of the 36th ACM Conference on Hypertext and Social Media (HT Adjunct ’25), September 15–19, 2025, Chicago, IL, USA. Pages 35–39.

Abstract

This paper explores the prospect of creating engaging user experiences and collecting leads through an interactive and gamified platform. We introduce Vocalize, an end-to-end system for increasing user engagement and lead acquisition through gamified voice competitions. Using audio processing techniques and LLMs, we create engaging and interactive experiences that have the potential to reach a wide audience, foster brand recognition, and increase customer loyalty. We describe the system from a technical standpoint and report results from launching Vocalize at 4 different live events. Our user study shows that Vocalize is capable of generating significant user engagement, which shows potential for gamified audio campaigns in marketing and similar verticals.

Keywords

lead acquisition, user engagement, gamification, generative AI, speech processing

1 Introduction

The active interaction of users with digital content has become a central pillar of success in modern business [1]. Active user engagement results in increased customer loyalty for existing clients [15], as well as the acquisition of new leads, often of high quality. However, generating user engagement through experiences that are both engaging and organic can be a challenging task for many organizations. One of the main hurdles in creating engagement, both for existing and new customers, is the lack of motivation and personalization when interacting with content [7]. A possible solution to these problems comes in the form of user participation through gamified competitions.

Figure 1: Screenshot of a Vocalize campaign interface at WeAreDevelopers 2024.

Figure 1. Example of a Vocalize campaign launched at WeAreDevelopers 2024.

Voice-driven chatbot interfaces, when combined with gamification concepts, can transform how businesses acquire leads and improve brand awareness. Such an approach makes interactions more engaging, and ensures personalized experiences that effectively capture attention and build lasting customer relationships [2]. As demonstrated in prior research, gamified conversational agents have shown potential to boost motivation in educational contexts [4], while voice-based agents have proven effective in crafting immersive narratives [14]. Furthermore, frameworks leveraging dynamic language adaptation and conversational cues have been shown to foster emotional engagement and trust, maintaining user-centered interactions [3, 10].

Contribution. In this work we present Vocalize1, an end-to-end system for increasing user engagement and lead acquisition through gamified voice competitions. Vocalize offers a multimodal interface that allows users to communicate with its underlying chatbot using their voice and text. By combining audio-based gamification and generative AI-based content personalization, we show how to improve lead acquisition and increase customer engagement, which in turn can result in greater brand awareness [13]. We support our claims by analyzing the results obtained in lead acquisition and user engagement from 4 venues at which we launched Vocalize-supported campaigns.

Figure 2: The underlying components of Vocalize. User audio inputs via WhatsApp are processed through keyword and shape scoring modules, with results enhanced by a generative AI and stored in a scoring database for feedback and leaderboard updates.

Figure 2. The underlying components of Vocalize. User audio inputs via WhatsApp are processed through keyword and shape scoring modules, with results enhanced by a generative AI and stored in a scoring database for feedback and leaderboard updates.

2 Vocalize

For simplicity, we will refer to end users who interact with Vocalize as users, while the term brand is used for those designing and launching specific Vocalize campaigns. The main purpose of Vocalize is simple; enable brands to acquire leads of higher quality and have their users more engaged by providing them a gamified experience, with brands optionally handing out prizes to best ranked users.

Vocalize is built on top of WhatsApp2 as its underlying channel for communication. WhatsApp has been chosen as it is the most popular messaging application in terms of monthly active users - more than 3 billion [12]. To do this, we utilize Infobip’s Answers platform3 for integrating with WhatsApp as well as to access OpenAI’s commercially available LLMs [8] to provide generative AI responses.

Competition outline. To initialize a competition, brands are only required to come up with a textual catch phrase and a desired image. For example, for a panoramic image of a city it could be the city skyline, or for an object it could be the edges of the object. Once users start interacting with the conversational agent over WhatsApp, they are provided with the predefined catch phrase and the outline of the image. Their task is then to record an audio message in which they repeat the catch phrase, but with an additional twist: the shape of the recorded audio message must also match the outline as close as possible. An example can be seen in Figure 1 where the catch phrase "I love Berlin" was used in addition to the image that represents the skyline of Berlin. Here we further refer to the outline of that image as contour. Figure 1 also shows how a user can improve their rank by recording an audio message that better matches the given contour. In summary, the main challenge of Vocalize in which users compete is to record an audio message that matches the campaign catch phrase in terms of content (spoken words), while also trying to match the shape of their audio recording (as visible in WhatsApp) with the target contour.

System architecture. A complete overview of Vocalize is provided by Figure 2. The WhatsApp chatbot interface enables users to communicate with Vocalize using natural language. If a user decides to participate in Vocalize’s main challenge and sends an audio recording, it is forwarded to Vocalize’s internal signal processing engine. The signal processing engine uses two separate components to formulate the score for the given user’s audio recording: a keyword scoring module and a shape scoring module. The user’s score is then logged to the database and the user is provided with feedback information. An example of such an interaction is shown in Figure 3.

Table 1. Statistics of the live events in 2024 where the user study was conducted. Larger events like WeAreDevelopers and Web Summit attracted significantly more voice recordings and longer interaction durations, while all events mostly demonstrated consistent message durations and engagement formats.

Event

Date

Target image

Target phrase

Voice recordings

Total duration

Median duration

WeAreDevelopers

17.07. - 19.07.

Skyline of Berlin

I love Berlin

6,321

4h 41m

2.25s

KulenDayz

30.08. - 01.09.

Skyline of Osijek

I love Kulen

1,257

46m

2.05s

GOTO Chicago

21.10. - 22.10.

Skyline of Chicago

Go to Infobip

1,216

1h 42m

3.38s

Web Summit

11.11. - 14.11.

Skyline of Lisbon

I love Lisbon

3,662

2h 22m

2.69s

2.1 Chatbot Interface

Figure 3: Users can interact with the system using natural language to receive scores, feedback, and personalized guidance in real time.

Figure 3. Users can interact with the system using natural language to receive scores, feedback, and personalized guidance in real time.

A typical way brands promote Vocalize is by advertising a QR code that initiates a conversation with Vocalize. This is useful for both virtual promotions, but also live venues as brands can print out the QR code on booths, posters or merchandise.

Once an interaction has been initiated, the conversational agent explains the rules of the game and collects contact information. As seen in Figure 3, the generative AI model enables the user to interact with the system using natural language. For example, the user can ask the chatbot about their score or about the prizes without having to click through menus. If the user sends an audio message, it is automatically forwarded to Vocalize’s underlying signal processing engine which assigns each recording a score. The score is returned, along with relevant metadata, e.g.: user’s position on leaderboard, number of attempts made, next best score, etc. The generative AI module is used to polish the metadata and present it to the user in a more natural form that mimics a real conversation.

2.2 Signal Processing Engine

Vocalize relies on an underlying signal processing engine to assign scores to user recordings. It consists of two components, both of which produce an individual score for any given audio: keyword score and shape score. A Vocalize campaign can have one or both components active for scoring, which is defined by the brand.

Keyword scoring module. This module rates how well the user’s spoken content matches the reference catch phrase (i.e., a target phrase). The module outputs a score ranging from 0 (total mismatch) to 1 (perfect match). To determine what the user said, we obtain a transcript of the user’s audio recording via automatic speech recognition. Depending on the campaign and brand requirements such as legal or privacy related considerations, we utilize an open-source solution, such as Whisper [9] or opt for a proprietary speech-to-text provider. The score itself is then computed by slightly modifying the Levenshtein algorithm for computing the distance between two strings [5]:

where D is the Levenshtein distance between the two strings and L is the length of the longer of the two strings (target phrase and user transcript). If the two strings do not share characters in common, this score will equal 0, while if they are identical, the score will be 1, which directly matches our requirements.


Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime