[Amiga][Down]

[ german ] [ english ]
[Scene] {}
 
[Special] 

| MP3 |
 
 

[Cover]
[Editorial]
[Contents]
[News]
[Hardware]
[Software]
[Workshop]
[Games]
[Special]
[Feedback]
[Etc]


MP3 - Basics

 
MP3 has become the accepted compression standard for music on almost every platform. But only a few know the excessive inefficiencies that are hidden behind the functionality. It is like the telephone, everybody uses it but hardly anyone know how it works.

 
The goal of this article is to bring the secrets of MP3 out into the open.

First of all; MP3 stands specifically for MPEG 1 layer 3. But as PC users know due to the problem with long file name extensions these are always shown as *.MP3. (Windows recognises file types based solely on their endings. Therefore a *.doc file is some kind of MS Word document, a *.xls an Excel file, a *.mpg is an MPEG, and a *.puke is a puke file :-)

 

Who created this?

 
Development for MP3 has been carried out by the Frauenhofer-Institute in Erlangen in cooperation with TU-Erlangen for the so called EUREKA-Project since 1987 (Not as new a technology as you might have thought?). Dr. Karlheinz Brandenburg (Phd in engineering) who at the time was director of the Department of Audio technology and Multimedia at Fraunhofer Institute is credited as the father of the MP3 process http://www.iis.fhg.de/amm

The goal of Eureka was among other things to drastically reduce the relatively high data rates of an Audio-PCM-Signal which can, for example, in a CD exceed 180 KBytes/s. With the previous nonlossy compression algorithms this could not be realized, therefore concentrated effort was made to develop a lossy compression process that would retain the highest possible quality. From this inspiration the MPEG coding process for audio and video was developed which became known in audio circles as MPEG 1 Layer 1.

One of the commercial uses for MPEG was the digital compact cassette (DCC) from Phillips. Thanks to MP1, delivered files were only a quarter of the size of the source data. The MiniDisc from Sony later employed a similar non conforming MPEG process with a compression factor of 5:1.

 

The further development

 
Soon afterwards the standards for MPEG layer 1 and 2 were differentiated. Layer 2, with its better compression, being the winner and only being sent into retirement by the follow on process MPEG 1 layer 3. The new concept utilized in MP3 was a decision to concentrate upon the maximum obtainable compression factor instead of upon the available CPU processing capability.


Layer and data rates

CD 1:1          180 kB/s      (stereo, 16 Bit. 44.1 kHz)
L1 1:4          48 kB/s       (stereo, 16 Bit. 44.1 kHz)
L2 1:6 - 1:8    32 - 24 kB/s  (stereo, 16 Bit. 44.1 kHz)
L3 1:10 - 1:12  16 - 14 kB/s  (stereo, 16 Bit. 44.1 kHz)

 
 
The newest technology is MPEG-2 AAC (Advanced Audio Encoding) and is a sequential further development of MP3. AAC offers, in addition to significantly better coding efficiency, (the same audio quality with a higher compression rate) extra functional capabilities such as the support for more programs with up to 48 channels or reception frequencies from 8 to 96 kHz and much more. Unlike MP3 free unlicensed AAC decoders will not be tolerated. Therefore this process will not be particularly interesting for the Amiga (unless of course Amiga Inc. licenses the process and builds it directly into the NG Amiga OS.)

 

How does MPEG function now?

 
The basis for all types of Audio coding is the psychology of acoustics.

This scientific field concerns itself above all else with physical measuring methods to determine at what point differences can not be perceived. Decision measures utilize the so called masking effect, the principal of which appears in two forms. For people it is not possible to resolve two neighboring tones of slightly different frequencies that also differ in volume as the softer tone will not be perceived and therefore the softer tone need not be retained.

[MP3 Simul Mask]
 
see also enlarged Figure MP3-simul_Mask-big.png

The bottom green curve shows, what sound pressure (SPL) is required to just hear a specific frequency. A tone of 1 kHz at 80 dB alters the minimum threshold drastically. The curve marked with 80 shows the area of mask. A tone of 2 kHz must now be louder than 50 dB instead of 10 dB to be heard next to the louder 1 kHz tone. A softer tone with a similar frequency to a louder tone is over a specific range masked by the louder tone to such an extent that the softer tone is unnecessary and can be removed. This is precisely how an MPEG-encoder functions. The resolution (of the sound differentiation) is dynamically processed to remove the portions of the signal that can not be perceived. The capability of the process is dependent upon splitting the entire frequency spectrum into 32 frequency bands (Layer 1 and 2). In layer three each of these 32 bands are further divided into 18 sub-bands. The encoder begins to function when the entire information content can not be carried within a preset bit rate. A 1 kHz sine wave with a bit rate of 128 kBits/s will for practical purposes be unaltered because all of the information will fit into the 128 kBit/s data stream. As one raises the frequency the information content also climbs and thereby also the amount of the required reduction (compression). Dependent on the quality of the applied algorithm and the availabilty of adequate CPU processing power an encoder can manipulate and optimize the incoming signal. The highest quality is delivered, naturally with the Fraunhofer encoder, because it works with signals upto 20 kHz, while the Xing encoder is limited to 16 kHz and is relatively noiser over the remaining frequencies, however it is also approximately twice as fast!

In addition there is a timely cover effect which results in quiter passages immediately before and after a sound louder than the average sound not being recognized. Hearing requires a recovery period from quieter as well as louder sounds until fully restored.

Thus the encoder has two important principles of functions for dynamic signals available in order to differentiate on a digital level between unrecognizable and recognizable signals.

The principles of masking so far described are sufficient to push the bitrate to 32 kbyte/s. This results in a compression ratio of 5:1 for CD sound. Further procedures are necessary for stronger reductions of the data rate to 16 kbyte/s. As this bitrate is normally used in stereo sound the next sequential step is self explanatory: an (imperceiveable) reduction in sereo information. In principle the impression of stereo is comprised of phase - and tone differences between the left and right channel. The human determines the direction of any sound event according to it.s freaquency.. Deeper sounds cannot be located (subwoofer-principle), medium sounds are located by the left or right ear via combination of loudness differences and timing of arrival; high frequencies (pitches) however are only located by the differences in loudness.

The Fraunhofer Encoder uses at first one of the basic functions of stereo techniques below 32 kbyte/s known as "Joint Stereo Coding", which is he middle/side-coding. The stereo signal is thereby taken apart into a middle (l+r) plus a side signal (L-R). As this side signal contains less information it results in reduced data compared to two fully separate stereo-channels. In this way large amounts of data are avoided as well as the reduction of stretches of irrelevant information. When using even lower than 8 kbyte/s data rates the "Intensity Stereo Coding" can be applied. For each frequency band only a sum channel and information regarding directions will be translated for a given frequency band. Because of the low data rates good quality results are obtained, however there exists noticable deviations from the original.

 

What happens inside the Encoder

 
At first the audio data is routed thru filterbanks. This results in partial data streams which are then assigned to the available bits in the quantisizer and coder ( quantisize is defined as the difference in resolution of the frequency band according to its requirements). The resolution varies dynamically between 2 and 15 bit ( originally 16 bit). A type of masking data base containing all possible maskings controls the assignment of bit resolutions.

This data bank is set up using data from very detailed test procedures with test subjects known to have excellent hearing. So far it had not been possible to measure human hearing and its realization using classical measurement techniques. Because of the progress made in digital techniques and DSPs we now have, at least initially, a very promising basis to support the development and improvement of the Encoder.

The search for the best possible type of masking for a given signal is very time consuming and largely responsible for the time required for the encoding. The logic of recogniation isutilized in an attempt to optimize this search via step by step aproximations (iteration).

Last not least the Huffman-coding as used for instance in package programs, is responsble for the suppression of redundant bits within the data stream (fi.e. bits without or with the same contents of information)

 

It's easy for the Encoder

 
In comparison, decoding is relatively simple. Although filtering of the original signal has to be recalled, a search for the matching masking method does not take place at all. Decoding in todays computer systems usually runs in the backgound without a burden on the CPU (except if you have a lame 68k proccssor). In addition the decoder has firm definitions, improvements of the procedure thus only happen with the encoder.

 

Encoding as fast as the wind

 
In the PC world the available selections from various vendors leaves no wish unfulfilled. However on the Amiga the selection is more limited. At this time I only know of three usuable encoders, among these are Lamer and Encoder. The latter is commercially available from Titan-software and should offer similar quality as the Fraunhofer Encoder. This account is not necessarily complete and is also not the actual topic of this article. We will deliver a test report as soon as the first commercial encoders are available.

 

Competition - MS Audio

 
As usual Micro$oft also attempts to conquer this market and guide it into a monopoly headed by Microsoft. By hitting the advertisement drum and by turning on the marketing machinery the lurching giant from Redmond stomps towards a frightened end user and threatens with raised finger: "Our audio standard is far better than MP3 and we are also faster at the coding process". It is natural to raise the question how they accomplish this within such a short time frame As we had mentioned before, the optimization of the coding procedure is a very time consuming matter and as such one has to think whether the patents.........

The results of the Fraunhofer Institute have been thoroughly documented and published - it is here where foreign knowledge was borrowed and (ab) used extensively. Of interest as well is that Micro$oft has so far not registered any patents.

As noted by different parties the quality of MS Audio was not really overly impressive compared to MP3, its speed is about that of the Xing Encoder, however with far poorer results.

 

The End

 
We don't need to be afraid of these matters , it is not expected that M$ will transfer its MS Audio to the Amiga. MP3 has its recognition and will be enjoyed for a long time by its users (it allows you to put 10 hours of music on one CD).

Robert Niessner     [ german ]    
Tom Lively     [ english ]    


- http://www.mp3.de
- http://www.mp3.com/faq
- http://www.mp3-online.de
- http://mp3.free-web.de/mp3/technik.html
- http://www.iis.fhg.de/amm/index.html
 
[Up]   {}