Skip to main content

Why 12 Notes?


By Eamon McGinn and Sergey Alexeev

MIDI note number 60 is middle C. Number 61 is C-sharp, 62 is D, and 72 is C an octave higher. The ordinary MIDI note map mirrors the piano keyboard, carrying centuries of Western musical practice into the digital world.

Twelve notes within the octave is so familiar that it rarely prompts much thought. But frequencies are continuous, and nothing in the physics says an octave must be divided into twelve pieces. So why twelve?

Our published paper in the Journal of Cultural Economics asks that question. We are interested in why pitch systems grow richer but don’t expand all the way to the continuum, and why twelve became so durable in Western music. We think economics has something useful to say about it.

An aesthetic choice is still a choice

The natural objection is that scales are an aesthetic matter, while economics is about money. But economics is really about trade-offs, and aesthetic choices are all about trade-offs. A painter with a limited palette, a poet working in a fixed metre, a composer working with a fixed set of notes: each gains something from the constraint and gives something up. Our paper asks what is being traded against what.

There are two sides to the ledger.

The benefit: more notes, more possibilities

Every pitch added to a scale forms new intervals with the pitches already there, opening new chords, melodic paths, and ways of building and releasing tension. The intervals listeners tend to find most consonant are associated with simple frequency ratios: the octave at 2:1, the fifth at 3:2, the fourth at 4:3, then the thirds. A scale built by stacking fifths captures many of these relationships early.

There is an important point here. Our acoustic calibration does not find that the total benefit of each extra note shrinks. Across the scale sizes we calculate, total harmonic value rises faster than linearly, and its fitted marginal value is increasing. What declines is harmonic value per possible interval pair, with later additions bringing in pairs of lower average consonance. The palette keeps getting richer, but not every new relationship is as valuable as the early ones.

The cost: every note comes with a bill

Historically that bill was physical: the cost of another pipe, string, hole, key, fret or mechanism. Instrument history gives a suggestive example. The flute achieved practical chromaticism with a single key in the 1670s. The clarinet, invented around 1700 and acoustically more demanding across registers, did not achieve full practical chromaticism until Müller’s 13-key system of 1812. This comparison is not a controlled test, but the sequence is consistent with engineering difficulty slowing the arrival of additional playable notes.

Today, with electronic synthesizers, the marginal physical cost of generating another frequency can be negligible. Yet performers must still learn where a pitch sits and how it relates to every other pitch. Composers, teachers, notation systems, controllers, software and other musicians must accommodate it. Adding a thirteenth pitch category to twelve creates twelve new pairwise intervals, before the extra chords and voicings are counted.

Where the two sides meet

Put the two sides together and something interesting happens. Falling construction costs make larger pitch systems more attractive. But the benefits do not simply “flatten out” in our calibration. The model produces a finite optimum when learning costs eventually rise more steeply than total harmonic value. The palette can expand as instruments become cheaper and still remain bounded when the physical cost of another pitch approaches zero.

The model’s result is this relationship, not the precise number twelve. It does not prove that twelve is uniquely correct, universal or historically inevitable. It shows how a musical system can expand and still have a place where it will stop, and how a twelve-note system can be understood as a trade-off among harmonic possibilities, technology, learning and coordination rather than as a law of acoustics.

Testing the idea against five centuries of music

We assembled 623 MIDI transcriptions of Western classical compositions by 48 composers from 12 countries, spanning 1485 to 1963. For each piece we counted how many of the twelve encoded semitone classes appeared, how evenly they were used, and how much the music drew on dissonant intervals.

Before 1900, encoded-class coverage generally rose. In the Renaissance and early Baroque, many works use only a subset of the available notes. Through the eighteenth and nineteenth centuries the count rises steadily as composers reach for more of the chromatic palette. After 1900, works in our sample used an average of 11.8 of the twelve classes, with no statistically clear continued increase. Within a measure capped at twelve, there was almost nowhere left to go.

Scatterplot showing unique pitch classes by year for compositions from 1485–1899 (red) and 1900–1963 (blue), with trend lines, a chromatic maximum at 12, and a diatonic level at 7 marked by horizontal lines.

Figure 1. Across 623 MIDI transcriptions of Western classical works, the number of distinct encoded semitone classes used in a piece generally rose before 1900 and then levelled off near twelve. Twelve is the maximum in this analysis because MIDI note numbers were grouped into twelve semitone classes. The chart shows saturation in chromatic coverage within that representation, not proof that historical pitch systems themselves contained exactly twelve categories. Source: McGinn and Alexeev (2026), Figure 6 (CC BY 4.0).

It’s important to acknowledge what the analysis cannot show. It cannot show that historical composers literally thought in exactly twelve pitch categories, reveal finer distinctions in performance, or establish twelve as a universal optimum. Surviving, digitized and transcribed classical works are not a random history of music either. The evidence concerns chromatic pitch use inside a modern twelve-class representation, not historical scale-system size itself.

Music did not stand still once encoded coverage was near its ceiling. Our dissonance measure continued to rise after 1900, and rose more quickly in the post-1900 fit. The evenness with which composers used the available classes could keep changing as well. Innovation shifted from expanding the encoded palette towards using a fixed palette more intensively.

A fixed palette is not the same as fixed music.

What MIDI 2.0 opens up

The ordinary MIDI 1.0 note map reflects the inherited semitone grid, but MIDI 1.0 never made microtonal music impossible:

  • Pitch Bend could move notes away from their nominal pitch.
  • The MIDI Tuning Standard allowed instruments to share user-defined tunings and retune individual notes.
  • MPE later gave each sounding note its own MIDI 1.0 channel, so Pitch Bend and other channel-wide messages could be applied note by note.

Those routes were real, but possibility is not the same as convenience. They could require tuning messages, channel management, extra configuration, compatible implementations, or careful control of the whole signal chain.

MIDI 2.0 extends rather than replaces MIDI 1.0. Its Channel Voice messages offer much higher resolution and stronger per-note mechanisms, including Note On pitch attributes and per-note pitch control. When implemented end to end, these features make it more direct to communicate a precise initial pitch and shape it independently, reducing reliance on channel-per-note allocation or separate tuning workarounds.

But a protocol can remove only part of the bill. Where does an unfamiliar pitch sit on a controller? How is it shown in a piano roll or notation system? Will it survive the journey from controller to DAW to plug-in to file to another musician’s setup? Who teaches the intervals, fingerings and repertoire? A tuning that can be generated but cannot be easily entered, seen, shared, learned or played with others is technically available and still costly to use.

MIDI’s great achievement has always been coordination: products from competing companies agreeing on a language. Twelve is not only a collection of frequencies. It is an ecosystem of instruments, interfaces, notation, teaching, repertoire, defaults and expectations. A standard can constrain the palette while also allowing strangers, software and instruments to coordinate. A new pitch language becomes more valuable when other people can speak it.

MIDI 2.0 therefore lowers an important technical barrier without deciding what musicians will do once that barrier falls. Will finer pitch control be used mainly for expressive intonation, bends and non-Western tunings within familiar structures? Will controllers, visualizations, notation and teaching evolve to make unfamiliar intervals easier to navigate together?

The model cannot answer those questions in advance, and MIDI 2.0 is not a controlled test of it. But it gives us a revealing real-world case: what happens when the cost of communicating pitch falls while the costs of learning, interface design and coordination remain?

The old problem was often how to build an instrument that could produce another pitch reliably. The emerging problem is how to make a broader pitch world easy to navigate together.


About the research

This article draws on Eamon McGinn and Sergey Alexeev (2026), “Why twelve notes? An economic model of scale size and evidence on chromatic pitch-class use in Western classical music, 1485–1963,” Journal of Cultural Economics. doi:10.1007/s10824-026-09607-y

Editor Note: There are many interesting projects from The MIDI Innovation Awards that explore ways musicians can achieve finer control of pitch.

2026

2025

2024

Earlier years (2021–2023, exact year not recorded on the site)