Back to Home

New Codec2 700C codec compresses speech up to 700 bps

FreeDV · 700C · MELP · Codec2 · 700 bps · ham radio · digital radio · single band modulation

New Codec2 700C codec compresses speech up to 700 bps


    In the program for digital radio transmission FreeDV, it will soon be possible to check the new codec in the work

    The author of the free voice codec Codec2, designed for super-dense coding of speech on voice channels, has released a new version of Codec2 700C in which he was able to encode distinguishable human speech at only 700 bit / s. This means that a three-second voice transmission with distinguishable speech will take only 260 bytes.

    Of course, such technologies are completely inappropriate for compressing music or other multimedia content, but this is an indispensable thing for communication in the face of severe restrictions on the bandwidth of communication channels. For example, when digitally transmitting sound from Mars.

    Such superdense compression can be useful not only for space applications, but also for amateur radio and for various military tasks, satellite communications, and encrypted devices. For example, now the U.S. Army uses the MELP (Mixed Excitation Linear Prediction) encoding standard , but it is the intellectual property of Texas Instruments (the algorithm and source code of the MELP codec at 2400 bps), Microsoft (transcoder 1200 bps), Thales Group (600 bit / s) and AT&T (noise preprocessor). The same proprietary MELP standard is used in satellite communications, secure voice communications and secure radio transmitters. The standardization and development of MELP was supported by the NSA and NASA.

    Obviously, people need a codec for a similar purpose, but free from patent encumbrances of MELP.

    Developer audio codec Codec2 - David Rowe (David Rowe). He has been leading the project for several years. The first alpha version of Codec2 was released back in September 2010 . This engineer previously had a hand in creating a free audio format for Speex speech coding , the development of which was discontinued in favor of the free Opus format . Then David set the task to achieve voice transmission in communication quality in a stream of 2400 bps and below, that is, to make a free alternative to MELP.

    “I continue to work on the development of a digital voice coding mode that can compete with single-band modulation ,” writesDavid Rowie - For most of 2016, I distracted from this work and was engaged in a paid project for a commercial high-frequency (HF) modem. But since December I’ve been working on a 700 bps codec again. The goal is to ensure that the quality is about the same as the current 1300 bps mode. This can be used in a coherent PSK modem, or maybe in a 4FSK modem when tested on HF channels. ”

    It is appropriate to clarify that the PSK modem is a device for a relatively new digital form of information transmission with narrowband on-off phase modulation.

    The author has done a lot of work to optimize the codec. The signal processing flowchart in the new codec is shown below. David Rowey writes that the key step in this algorithm is resampling (resampling), when the time-varying number of harmonics amplitudes is converted to a fixed number (K ​​= 20) of samples. At lower frequencies, more samples are taken than at higher frequencies, which corresponds to the logarithmic perception of the human ear. Experimentally, David came to the value of K = 20.



    The main part of David Rowey’s work concerns perceptual compression of sound in such a way as to best match the logarithmic characteristics of human hearing, when the susceptibility to different frequencies varies according to the logarithmic law.


    3D graph of the ratio of the amplitude in dB over time (300 frames) with the oversampling parameter K = 20 frequency vectors for the sound sample hts1a (you can listen to it in the table below). It can be seen that the signal changes over time and low values ​​at high frequencies that are worse perceived by the human ear.

    In general, Bruce Perens raised the problem of the shortage of free codecs in the range up to 5 kbit / s in 2009. He contacted Speex developers and invited them to study the situation. Codec2 is based on the scientific work of the 60-80s and does not seem to fall under current patents. Sinusoidal coding of speech was first mentioned in 1984, and Rowe described in detail the techniques of harmonic sinusoidal coding in his 1997 research paper. The codec is published under the free license of LGPL2.

    In the samples below, you can compare the samples of the previous version of the codec at 1300 bps and the new version at 700 bps.
    Sample1300700C
    hts1aListenListen
    hts2aListenListen
    forigListenListen
    ve9qrp_10sListenListen
    mmt1ListenListen
    vk5qiListenListen
    vk5qi 1% BERListenListen
    cq_refListenListen
    Each person has their own hearing characteristics, so the author asks to leave feedback: how distinct are 700 bit / s samples compared to 1300 bit / s? David Rowey believes that they are about the same: some samples are slightly better (cq_ref), while others are slightly worse (ve9qrp_10s, mmt1). Artifacts are different everywhere. But in any case - this is an almost twofold reduction in the band!

    For comparison, here is a comparison of the alpha version of the Codec2 v0.1 codec (2550 bps) from 2010 and the proprietary MELP codec (2400 bps).

    Male voice:
    Original
    Codec2 v0.1 (2550 bps)
    MELP (2400 bps)

    Female voice:
    Original
    Codec2 v0.1 (2550 bps)
    MELP (2400 bps)

    In the coming weeks, David Rowey is going to open the 700C codec through the interfaces for the FreeDV digital radio program and conduct the first tests on the air.

    Read Next