Signal & Syntax · Guide
GSM-7, UCS-2 and SMS Segments Explained
One curly apostrophe turns a 160-character SMS into a 70-character SMS. The rules below explain why, and how to keep a message in one segment.
5 minUpdated ScoreMachine team
One SMS Is 140 Bytes
The SMS payload is 140 bytes, set in 3GPP TS 23.040. The 140 bytes hold 1,120 bits. The encoding decides how many characters fit in those bits.
160 × 7 bits and 70 × 16 bits both come to 1,120 bits. The 160 and the 70 are the same space, measured in two alphabets. Operators bill per SMS, so the byte limit is also a price limit.
Every character in a message uses the same encoding. A message has no way to mix the two. One character outside the smaller alphabet sets the encoding for the whole message.
The limit dates from the first GSM standards and still holds on 4G and 5G networks.
The GSM-7 Alphabet
GSM-7 is the default SMS alphabet, defined in 3GPP TS 23.038 (formerly GSM 03.38). The basic set holds 128 characters. Latin letters, digits, common punctuation and a set of accented letters sit in the basic set. Each basic character takes one 7-bit unit. Digits, spaces and line breaks each take one unit too. The basic set also carries Greek capitals such as Δ, Φ and Ω.
The extension set adds ^ { } \ [ ~ ] | and €. Each extension character takes two units: an escape code, then the character.
Example
Basic set, 1 unit: A–Z a–z 0–9 @ £ $ ¥ è é ù ì ò Ç Ø ø Å å Ä Ö Ñ Ü § ¿ ä ö ñ ü à Extension set, 2 units: ^ { } \ [ ~ ] | €
A 160-character message with one € counts 161 units. The message then splits into two segments.
The basic set has no curly quotes, no en or em dash and no ellipsis character. Each of those arrives from word processors, phone keyboards and copied templates.
3GPP also defines national shift tables, such as Turkish and Portuguese, for extra characters. Support depends on the operator and the handset. A sender cannot rely on a shift table.
When a Message Switches to UCS-2
UCS-2 is the 16-bit encoding for characters outside GSM-7. One character outside GSM-7 switches the whole message to UCS-2. The limit then drops from 160 characters to 70.
The characters behind the switch fall into five groups:
- —Curly quotes and apostrophes: ’ ‘ “ ”
- —En and em dashes: – —
- —The ellipsis character: …
- —Emoji
- —Non-Latin scripts, such as Arabic, Chinese, Cyrillic and Devanagari
Word processors and phone keyboards insert curly quotes on their own. A template pasted from a document picks them up without anyone typing one. The sender sees a normal message, and the operator bills a UCS-2 message.
UCS-2 gives every character two bytes, plain letters included. Once one curly quote is present, each letter costs twice the space.
The switch also hits characters that look like GSM-7. A non-breaking space, copied from a web page, forces UCS-2 without showing on screen.
Emoji outside the Basic Multilingual Plane take two UCS-2 units each. A message with one such emoji and 69 other characters counts 71 units, and splits.
How Segments Work
A message over one segment's limit splits into segments. Each segment carries a user data header (UDH). The UDH tells the handset how to rejoin the parts. The UDH takes six bytes of the 140, and 134 bytes remain for text.
The handset shows the rejoined parts as one message. A two-segment message costs twice a one-segment message. Each segment is billed as one SMS. How A2P SMS reaches a handset covers the billing path. A 161-character GSM-7 message costs two segments. A 71-character UCS-2 message also costs two.
Two rules keep a character whole:
- —An extension character never splits across two segments.
- —A surrogate pair, such as an emoji, never splits either.
A split landing mid-character moves back one unit. The segment then holds 152 units, and the character opens the next segment.
The segment count climbs in steps. In GSM-7, 306 characters fit two segments and 459 fit three. In UCS-2, 134 characters fit two segments and 201 fit three. A 200-character GSM-7 message costs two segments. The same text with one curly quote costs three.
A Worked Example
The sample message below comes from the free SMS Segment Counter.
Example
We’ll call you at 10:30 tomorrow. Reply C to cancel.
The apostrophe in "We’ll" is U+2019, a curly apostrophe. That one character moves the message to UCS-2.
The message fits one segment both ways. The headroom changes from 18 characters to 108. A 19-character addition pushes the UCS-2 version into two segments. The GSM-7 version takes 108 more characters before a split. Swapping the apostrophe changes nothing a reader notices. The SMS Segment Counter flags the apostrophe on first paste.
Keeping a Message in One Segment
Five habits keep a message in one segment:
- 01Replace curly quotes and apostrophes with straight ones.
- 02Replace en and em dashes with a hyphen, and the ellipsis with three dots.
- 03Count each GSM-7 extension character as two.
- 04Fill merge fields with the longest value the field holds, then count.
- 05Count the final text before the send, not the draft.
Links count too. A 23-character short link takes 23 units of the 160. Merge fields change the count. A long first name pushes a 158-character template past 160. A curly apostrophe inside a customer name switches the whole message to UCS-2.
Shorten dates and times. 10:30 takes five units, and "half past ten" takes 13. Keep the brand name out of the body when the sender ID already carries the brand.
Test a template on the longest real record in the file, not on the sample row. A translated template changes both the length and the encoding.
Arabic, Chinese and Devanagari messages need UCS-2 by design. Write those templates to 70 characters, or 67 per segment when split.
The free SMS Segment Counter shows the encoding, the count and the segments as you type. The counter highlights each character forcing UCS-2 and offers a GSM-7 swap. The text stays in your browser.
Terms in this guide