Composing with and for machines
Musical Structure

Groupings are the basis for musical structures that delineate a work’s form and proportions. Smaller entities or ideas serve as the foundation on which larger structures are built. In music, these include the motif (motive), a short musical pattern or fragment that may be expressed in melody, harmony, or rhythm. While the size of a motif is variable, it is commonly regarded as the shortest subdivision of a theme or phrase that maintains its identity as an idea and is used to generate longer statements. A famous example of a motif occurs at the beginning of Beethoven’s Symphony No. 5 in C Minor, Op. 67 (Figure 5.1):

Figure 5.1 — Motif in Beethoven’s symphony no. 5

The essence of this motif is a pattern of three short elements of the same pitch followed by a longer element of a lower pitch. Listen to the opening movement of the work to hear how the motif is presented and transformed. Beethoven’s 5th is a fantastic example of the bounty of music that can be grown from a tiny seed.

Smaller elements, such as motifs, form larger organizations, such as sequences, phrases, melodies, and themes. A motif or other short unit that is repeated at different pitch levels creates a sequence. A musical phrase is analogous to a linguistic phrase: both group smaller elements into larger, cohesive statements that are bounded, representing complete thoughts. A phrase’s length can vary, but it is generally longer than a motif. Examples of musical phrases are found across genres and time periods, from the saxophone part of John Coltrane’s Giant Steps (Figure 5.2; notice how the video slows and stops at phrase boundaries):

Figure 5.2 — Animated notation of John Coltrane’s “Giant Steps” (Cohen, 2007).

to Beethoven’s 9th Symphony Ode To Joy (Figure 5.3), which Jesse Strickland discusses in this video:

Figure 5.3 — A musical phrase from Beethoven’s Symphony No. 9, Mvmt IV, “Ode to Joy” (Strickland, 2018).

Phrases are the basis for melody, which is a contoured statement with a definitive (and often singable) character, formed from the combination of pitch and rhythmic sequences. Note that not all phrases are melodic — rhythmic phrases can be produced on non-pitched percussion instruments. As with linguistic phrases, musical phrases occur across musical styles and cultures. The conventions of these styles and cultures, as well as personal aesthetic preferences and perceptual and cognitive capabilities, are factors in phrasing. A string of short statements can come across as choppy and basic, lacking flow and sophistication. Long statements without breaks inhibit the listener’s ability to find boundaries for perceptual groups and challenge the limits of short-term memory. Content is also important: phrases should progress in ways that make the logic and the message of the communication clear (both within phrases and between them). Listen to examples while thinking about the length, proportion, content, and the “logic” of musical phrases. From these analytical exercises, you can develop profiles of phrasing that you find interesting, which can then be the basis for algorithms that determine how a musical machine will generate phrases.

When a musical idea is presented, repeated, and developed to the point where it is salient, memorable, and identifies the work as a whole, it becomes a theme. Themes are typically larger, coherent statements made from combinations of motifs, phrases, melodies, and connective material. A theme is a central idea you take away from an artistic experience, whether you listen to music, read a book, watch a movie, or view a piece of art. For music-generating machines, crafting an idea that becomes thematic is a challenging compositional and programming problem. Some aspects of thematic writing, such as repetition and development, can be algorithmically deployed, though one must keep in mind that reiteration can arouse comforting familiarity and boredom, and that there is a fine line between the two. Determining what content is suitable to become a theme is also difficult. One approach is to analyze (manually or computationally) themes you like, are famous, or represent a desired style to determine defining characteristics that serve as the basis for the machine’s generative process. Alternatively, you may approach the process as an experimental composer, trying out new ideas, evaluating the results, and tweaking parameters toward something you find engaging.

Form

The highest structural level of a musical work is form, which describes the qualities, sequences, and relationships of the largest sections. When setting out to write a new work, you may start with a previously established form. The constraints in doing so can direct your attention away from worries about how many sections to write and how they will fit together as a whole. Pre-established forms have examples to emulate and diverge from as you see fit. When notating a form, the sections are represented by letters of the alphabet (e.g., “A”), which are delineated by various musical factors, including key, chord progression, melody, rhythm, and instrumentation. Contrasting sections are noted by different letters (e.g., A B). If a section is repeated but varied, it is represented with a superscript or prime label (e.g., A A1 or A A′). Sub-sections are indicated by lowercase letters (e.g., a, b). Some examples of form include:

  • One-part: a single section with limited musical ideas (usually one) that is developed over the work.
  • Strophic (A A . . . ): one primary section that is repeated many times, with variations in verses (usually sung) occurring at every repetition.
  • Binary (A B): both sections are of roughly equal length. Strong punctuation in the middle delineates the two, with the second section contrasting the first in terms of key, activity, and stability.
  • Ternary (A B A): the A and B sections are independent and usually able to stand on their own, contrasting in terms of theme, activity, harmony, register, and articulation.
  • Variation: a central idea is introduced and then is transformed in each subsequent section in terms of tempo, rhythm, duration, density, or harmony such that each variation has its own character. The final variation usually breaks from the theme more dramatically, is longer, and is more complex.
  • Rondo (A B A C A  . . . ): the periodic return of the A section provides a foundation from which to present contrasting themes in the B, C, and so on, sections.
  • Sonata form has particular significance in the history of European Art Music. It is the frame for a metaphorical discussion involving ideas, conflicts, dialogue, and resolution. The first section is the exposition, within which two conflicting ideas are presented. These ideas are transformed and juxtaposed against new material in the development, which is less stable and discursive. Material from the exposition is represented in the third section, the recapitulation, though with some kind of resolution (e.g., in harmony, where both themes appear in the tonic key).
  • The foundation of blues form is a I–IV–V chord progression that often occurs over 12 bars. The sequence of the chords can vary, but it typically is I (mm. 1–4), IV (mm. 5–6), I (mm. 7–8), and V (m. 9). Measures 10–12 vary using different combinations of V, IV, and I. This form has been enormously influential in other musical styles such as rock.
  • Pop form is ubiquitously present in a variety of musical styles. The basic ingredients are intro—verse—chorus—bridge—outro with a dizzying number of permutations that include repetitions and omissions of these elements. The verse and chorus are the most important parts of pop form. The verse sets the stage in instrumentation, harmony, and rhythm and paves a path for the chorus, which is intentionally catchy and is often climactic. Intros, outros, and bridges provide context, anticipation, respites, expositions, and resolution.
  • Form can also be abstract. Start with the image of a tree. If the music began at the base of the trunk, how might it develop to ascend upward and stretch out among the branches, and settle in the leaves? This is but one of countless possibilities. Abstract forms have been the inspiration for graphic scores, which provide performers with interpretative power. These possibilities are intriguing for musical machines. Instead of prescribing exactly what notes the machine is to play, it could be given a map and a set of directions that could be used to navigate it.
Rhythm and Time

The primary units of rhythm are relative durations, measured from the onset of one note to the onset of another, known as an inter-onset interval (IOI). Musical time may unfold according to regular temporal intervals or not. The latter, known as free time, can be delineated infinitely, making machines’ ability to accurately measure and produce temporal intervals valuable in realizing new temporal structures. Time marked by regular intervals, known as measured time, features palpable, evenly spaced beats and is found in music across the globe. Music in measured time often contains multiple periodicities occurring simultaneously, some slower, some faster, in rates that are related by small-integer ratios (e.g., a kick drum playing every 500 ms and a hi-hat playing every 250 ms). These rates are subdivisions and concatenations of an intermediate isochronous (equally-spaced) periodicity known as the tactus, which connects and provides a reference for the rhythms present in a musical texture (this is the rate that you tap your foot to and that is typically notated on a score). The rate at which the tactus occurs is known as tempo.

Some beats are felt more strongly than others. When strong–weak beat sequences form a pattern, meter emerges, which can be described using a time signature such as 3/4. The combination of different periodicities and accent patterns forms a metric hierarchy (Lerdahl & Jackendoff, 1983), as seen in Figure 5.4. The y-axis represents the metrical level of the periodicities present (eighth, quarter, half, and whole-note durations), and the x-axis shows metrical positions, which are temporal locations measured in beats. Dots indicate the presence of a beat (or not) for a given metrical level at a particular metric position. For example, at the quarter note level, there are dots at positions 1, 2, 3, and 4, while at the whole-note level, there is a dot only at position 1. As position 1 has dots at every metrical level, it is the strongest, corresponding to our notion of the downbeat as the center of a metric cycle. The strength of other metric positions can be determined similarly by counting the number of dots present; for example, beat 4 is weaker than beat 3. Different meters have different patterns of strength for metric positions, depending on how many metrical levels are present.

A metrical hierarchy for 4/4 meter
Figure 5.4 — A metrical hierarchy for 4/4 meter

The length of a rhythmic sequence can expand or contract by adding or subtracting time from the beginning, end, or middle of the pattern. A similar idea can be applied to meter, where expansion is achieved by adding beats (e.g., 3/4 → 4/4) and contraction involves subtracting them (e.g., 4/4 → 3/4) as the piece progresses. This technique of metric expansion and contraction is used in Rise of a City (2009) and Tempo Mecho (2017), shown in Figure 5.5:

Expanding and contracting meter in Tempo Mecho (Barton, 2017)
Figure 5.5 — Expanding and contracting meter in Tempo Mecho (Barton, 2017)

Different beat values can be used; for example, a measure in 7/8 could expand to 4/4. These mathematical operations are fundamental ways to transform a musical sequence, creating a sense of movement and syncopation as they deviate from established patterns. They are discussed in more detail in Chapter 9 of the main text.

Manipulating a rhythm’s rate changes the length of constituent intervals and the durations over which they occur. Slowing (or lengthening) a rhythmic pattern while maintaining its intervallic proportions is known as augmentation while speeding it up (or shortening it) is known as diminution. At some rates, the character of the rhythm changes, a phenomenon explored in the context of rhythm production (Barton et al., 2017). At high speeds, the progression of notes is so fast that there is a point where individual articulations are no longer discernible. At this threshold (which is usually close to the rate at which periodic repetition is perceived as continuous sound: 20 Hz in human beings), the rhythm becomes texture and, in some cases, pitch. Varied textures can be created by speeding different rhythmic patterns beyond this threshold.

Rate changes can be immediate (terraced) or proceed according to a contour. Linear changes in rate are the most obvious (and are generally assumed with the directives accelerando or rallentando), but any mathematical function can create a contour that will govern the rate at which a machine percussionist will play.

In Western Art and popular music, rates typically relate to each other by 2n (a half note is played twice as fast as a whole note; a quarter note is played four times as fast as a whole note, etc.). Rates that change within a voice that are not mathematically related by 2n are traditionally notated using tuplets (Figure 5.6):

Tuplets
Figure 5.6 — Tuplets

Tuplets can be nested within each other to create even more changes in rate (Figure 5.7):

Nested tuplets
Figure 5.7 — Nested tuplets

Divisions and concatenations by three and its multiples are also common (but not as plentiful as two). As ratios between rates become more complex, they become rarer, with fractional or irrational relationships seldom seen (such are hard to play!).

Forming more complex rhythmic composites involves coordinating different rhythms so that simultaneities and displacements of individual notes occur. An example of rhythmic coordination between two voices is hocket, which involves distributing a rhythmic pattern between two voices in alternation. Hocket was a prominent device used in the vocal music of thirteenth- and fourteenth-century Europe, and its utility in creating rhythmic and timbral interest persists today. Figure 5.8 shows an example of a hocket where a rhythm in the first bar is distributed between two voices in the second bar:

Hocket
Figure 5.8 — Hocket

Polyrhythm is a general distinction that refers to the simultaneous sounding of multiple rhythms that contrast with each other in rate or length causing an uneven division of the prevailing meter. Polyrhythmic rates may be simple divisions or concatenations of the tactus, or their relationship may be more mathematically complex. One type of polyrhythm is a cross-rhythm, which refers to a rhythmic grouping that does not evenly divide the meter, thus “crossing” the barline, as seen in the top voice of Figure 5.9:

A cross-rhythm
Figure 5.9 — A cross-rhythm

This creates rhythmic configurations of different durations (lengths) that will synchronize, dislocate, and synchronize again, depending on the mathematical relationship between the two. Displacement is an essential part of the rhythmic relationship that results, as the beginning of one part aligns with different positions of another upon each repetition.

Polymeter occurs when two or more meters derived from the same beat are present simultaneously. As a result, two different accent patterns are infused in the same beat. For example, 3/4 and 4/4 are combined in Figure 5.10:

Polymeter
Figure 5.10 — Polymeter

The concepts of cross-rhythm and polymeter are similar and can be confused. The reason to use one or the other is partly based on notation. Writing two (or more) meters simultaneously may confuse some human performers when trying to coordinate with others (bar numbers will differ between the parts), so using a single meter with cross-rhythms can be more straightforward. Machines have no problem dealing with such complexities, so one may choose a polymetric representation to make the rhythmic groupings clear in a descriptive score. Cross-rhythm and polymeter may also sound different. If one is playing a rhythmic pattern lasting three quarter notes in 4/4, the implicit accent patterns of that meter remain, affecting points of emphasis in the performance of the rhythm. The same sequence played in 3/4 would have different qualities, given the different accent patterns of that meter. Human performers assume these interpretations; they need to be specified by the programmer when working with musical machines.

Temporal displacements can occur between two rhythms of the same length, creating a phase relationship. In a canon, multiple voices sound the same melody with each statement offset in time. The melodic repetition directs attention away from the pitch sequence and toward rhythmic and harmonic features. As the musical image is superimposed against itself, changes in its position create complex composite textures. Steve Reich was interested in a dynamic version of this musical process in which phase relationships changed continually, implementing it in works such as Come Out (1966), Piano Phase (1967), and Clapping Music (1972).

Rhythmic voices can play at different rates simultaneously, creating effects such as half- or double-time and polytempo. In half- or double-time, some voices slow down or speed up (e.g., the kick drum and snare), while others maintain their pace, creating a hybrid that suggests movement has changed without actually varying the tempo. Polytempo describes a musical texture perceived as having more than one tactus. Polymeter and polytempo can be combined for even greater complexity. In a polytempic canon, voices proceed at different rates, creating converging (the faster voice starts later than the slower voice; the two synchronize at a future point), diverging (voices start together, but the slower voice finishes later, as in Ligeti’s Poème Symphonique), converging-diverging (the faster voice starts later than the slower voice, the two synchronize and the faster voice finishes first), and diverging-converging forms (the voices start together, diverge due to the rate differences which are reversed at some point, allowing the voices to reach an eventual point of synchronization). Conlon Nancarrow explores these rhythmic possibilities in fascinating depth in his Studies for Player Piano (see Thomas, 2000).

Expressive timing, the minute temporal variation from ideal notated proportions, is a distinctive aspect of human performance. It frees performance from rigidity (a complaint of many quantized machinic performances), but not to the point where the rhythmic interpretation is incorrect. This temporal flexibility in part defines musical styles, from the rubato of a Chopin piano prelude to the feel of jazz drummers such as Elvin Jones. The character of expressive timing is complicated and depends on a number of factors, including motor strategies and structural interpretations (Clarke, 1985).

Pitch

Pitch is a psychological phenomenon that corresponds to the vibrational frequency of a sonic object. The class or chroma of a pitch refers to a quality that is shared between the “same” note in different octaves. For example, the pitch A4 (chroma A, octave 4) and the pitch an octave above it (A5) are the same in some meaningful way, thus they are labeled with the same symbol. Pitch height is independent of chroma; it is one of the ways that two pitches of the same chroma differ. It is a spatial metaphor that we use to distinguish the “distance” between pitches, where A5 is “higher” than A4. We measure the distance between pitches in intervals, which are determined by segmenting the octave evenly (e.g., into 12 semitones). The register of a piece or instrument refers to the vertical range of pitches that it contains or is capable of.

Musical organization involves combining different pitches vertically (e.g., in a chord) or horizontally (e.g., in a melody); these groups form the basis of harmony. The characteristics and boundaries of human memory contribute to harmony perception, which a machinic generator could model (or not). The space of available pitches is defined by a tuning system such as just intonation, where intervals are based on whole-number frequency ratios, or equal temperament, which divides the octave into equal intervals (this is the system that is commonly implemented in modern musical instruments). The less-explored world of microtonality exists between the discrete steps of the chromatic scale.

From the expansive pitch space afforded by a tuning, smaller groups of pitches (or pitch classes) such as sets, scales, and keys become the basis for the harmony of a musical section or work. The most general of these groups is the pitch set, which is a collection of musical pitches. Setting constraints on a group of pitches is a great way to liberate yourself from struggles with seemingly endless possibilities. When pitch sets are ordered and conventionalized, they are known as scales. The pitches of a scale are known as degrees, where the lowest is the first (also known as the tonic) and ascend/descend incrementally (e.g., in the scale C D E F G A B, E is the third scale degree). Scales are identified by the pattern of intervals between their pitches (e.g., the diatonic scale follows the pitch interval pattern W W ½ W W W ½, where W is a whole tone and ½ is a semitone). Scales typically comprise pitch classes rather than absolute pitches, which allows them to be repeated in different octaves. There are thousands of scales in the world. Common ones in Western music include the diatonic scale (from which major, minor, and the modes can be derived), chromatic, whole tone, pentatonic, and octatonic. Scales are the basis for keys, which form the harmonic foundation of a musical passage, section, or work. Harmony does not need to be static: modulations to new harmonic areas are a fundamental way to create interest. Elaborate systems using scales and keys have formed, such as the Major/minor tonal system developed during the Common Practice Period of European art music (and is still very much alive in the music of today), 12-tone serialism developed in the early part of the twentieth century, and modal harmony in 1960s jazz.

Texture

Texture can be described using formal concepts from music theory that describe the roles and relationships of voices, which may refer to sung or instrumental parts. Monophony consists of a single primary voice, for example, as in Gregorian Chant. A homophonic texture is one where a primary voice is supported by one or more additional parts that provide harmony and often rhythmic contrast. This texture is found in most pop songs, where a singer voices a melody that is accompanied by guitar, piano, bass, or drums that fill out the harmonic and rhythmic landscape. Heterophony features simultaneous variation of a single melodic line, as heard in Arabic classical music and Indonesian gamelan. Music that contains multiple, independent, simultaneous voices is polyphonic. Polyphony developed in the vocal music of Christian churches of Europe during the Medieval and Renaissance periods and became a foundation upon which the compositional practice of European Art Music was built.

References
  1. Barton, S., Getz, L., & Kubovy, M. (2017). Systematic variation in rhythm production as tempo changes. Music Perception: An Interdisciplinary Journal, 34(3), 303–312.
  2. Clarke, E. F. (1985). Some aspects of rhythm and expression in performances of Erik Satie’s “Gnossienne No. 5.” Music Perception, 2(3), 299–328.
  3. Cohen, D. (2007, January 3). Animated sheet music: “Giant Steps” by John Coltrane [Video]. YouTube. https://www.youtube.com/watch?v=2kotK9FNEYU
  4. Lerdahl, F., & Jackendoff, R. S. (1983). A generative theory of tonal music. The MIT Press.
  5. Strickland, J. (2018, October 25). Building music: Phrases and periods — Two minute music theory #35 [Video]. YouTube. https://www.youtube.com/watch?v=50fOnEkIcOY
  6. Thomas, M. E. (2000). Nancarrow’s canons: Projections of temporal and formal structures. Perspectives of New Music, 38(2), 106–133.