Audiolizr (audio and analyzer) is an API built and deployed with BentoML to transcribe Youtube videos and extract the following metadata:

This API can be used to provide a summary and additional information to understand any youtube video or audio content.

This service is deployed on AWS EC2 on a GPU-powered g4dn.xlarge instance. (see deployment section for details)


I've used audiolizr to process this (interesting) TEDx short video


Here a demo of audiolizr in Streamlit.


As shown in the following diagram, the video moves through different runners to

  1. download the audio
  2. transcribe the audio into text
  3. extract keywords
  4. extract named entities
  5. summarize the text

note: runners 1 and 2 are executed sequentially and runners 3, 4 and 5 are executed concurrently

Here's the JSON output that you'd get at the end of the pipeline:

  "transcript": "How much do you get paid? Don't answer that out loud. But put a number in your head. Now, how much do you think the person sitting next to you gets paid? It turns out that pay transparency, sharing salaries openly across a company, makes for a better workplace for both the employee and for the organization. You see, keeping salary secret leads to what economists call information asymmetry. This is a situation where in a negotiation, one party has loads more information than the other. And in hiring or promotion or annual raise discussions, an employer can use that secrecy to save a lot of money. Imagine how much better you could negotiate for a raise if you knew everybody's salary. Now, I realized that letting people know what you make might feel uncomfortable, but isn't it less uncomfortable than always wondering if you're being discriminated against, or if your wife or your daughter or your sister is being paid unfairly? Openness remains the best way to ensure fairness. And pay transparency does that.",
  "metadata": {
    "keywords": [
        "pay transparency",
        "paid unfairly",
        "information asymmetry",
        "call information asymmetry",
        "sharing salaries openly",
        "call information",
    "entities": [
        "entity_text": "one",
        "entity_label": "CARDINAL",
        "start": 438,
        "end": 441
        "entity_text": "annual",
        "entity_label": "DATE",
        "start": 521,
        "end": 527
    "summary": "If you know everybody's salary, you can save a lot of money. And in hiring or promotion or annual raise discussions, an employer can use that secrecy to save money. Pay transparency, sharing salaries openly across companies, makes for better workplaces for both the employee and for the organization. I realized that letting people know what you make might feel uncomfortable, but isn't it less uncomfortable than always wondering whether your wife or your daughter is being paid unfairly?"


Run locally

Run the following commands to start a fresh environment with the needed dependencies:

cd audiolizr/
pipenv install 
pipenv shell

# install whisper with pip
pip install git+https://github.com/openai/whisper.git
# install spacy language model
python -m spacy download en_core_web_md 

To serve the API locally, run the following command.

cd src/
bentoml serve service:svc --reload

To serve the API in production mode (and enable multiple api workers), run the following command (keep --api-workers low to avoid hammering the RAM)

cd src/
bentoml serve service:svc --production --api-workers 2

If everything works as expected, build the bento to prepare the deployment:

cd src/
bentoml build

Here's what you'll see when it's done:

Run the API from the Docker image (with port forward) and check that everything's running fine from the host

docker run -it --rm -p 3000:3000 speech_to_text_pipeline:m57a6etzlg4imhqa serve --production --api-workers 2

Head over http://localhost:3000 to try out the API

Deploy to EC2

To deploy the API on AWS, you need to follow these steps:

1. Install the required tools

2. Setup the deployment config

To deploy the BentoML service on AWS, we will use bentoctl, a tool that helps deploy any machine learning model on any cloud infrastructure.

bentoctl uses Terraform under the hood.

Since we're going to deploy the service on EC2, we first need to install the AWS EC2 operator to generate and apply the Terraform files

bentoctl operator install aws-ec2

The deployment configuration must be detailed in a deployment_config.yaml file (this one is contained in the bentoctl folder)

This file contains details about the instance type, the AMI ID, the region, etc.

api_version: v1
name: audiolizr-bentoml
  name: aws-ec2
template: terraform
  region: eu-west-3
  instance_type: g4dn.xlarge
  # points to Deep Learning AMI GPU PyTorch 1.12.0 (Ubuntu 20.04) 20220913 AMI
  ami_id: ami-0aed2a32a2ea85e67
  enable_gpus: true

3. Generate the Terraform files

Simply run

bentoctl generate -f deployment_config.yaml

This will generates the main.tf and bentoctl.tfvars

4. Build the Docker image and push it to ECR

The image upload may take time depending on your bandwith

bentoctl build -b speech_to_text_pipeline:dzunp5dzo2uhmhqa

5. Apply the Terraform file to deploy to AWS EC2

In this step, we apply the Terraform file to deploy the docker image on an EC2.

This step fires up an instance and configures all the needed installation

When everything is up and running (~ takes approximatively 10 minutes), you will be prompted with a your API URL.

🎉 !

6. Test the API

Once the API is deployed, you can try it out either from the browser or via a python client.

Let's query it from an ipython terminal and compare the inference time with my locally running smae API (with no GPU)

By upload an audo file

By downloading a youtube video

  1. Remove the service when you don't need it anymore
bentoctl destroy -f deployment_config.yaml