This page collects outside readings and test results on LTE throughput. Each section gives the source, its key numbers, and a short note that connects them to the LTE protocol stack or to a throughput test setup.
Keep the dates in mind while you read. The Ericsson article is from 2008 and the iPerf blog post is from 2011, as their links show. Some of the problems they describe have since been addressed in the 3GPP specifications, and the notes say where.
The page covers the following topics.
- TCP Performance Degradation of In-Sequence Delivery in LTE Link Layer
- Initial field performance measurements of LTE
- Using iPerf to Troubleshoot Speed/Throughput Issues
- Stochastic Forecasts Achieve High Throughput and Low Delay over Cellular Networks
- Reference
TCP Performance Degradation of In-Sequence Delivery in LTE Link Layer
Why can a link layer that recovers every lost packet still slow TCP down? This paper answers that with the RLC reordering timer. In-sequence delivery holds every packet behind a gap until HARQ recovers the gap or t-Reordering expires, and TCP sees that wait as a longer RTT. The paper is available as a pdf.
Hyun-Seo Park, Jae-Yong Lee and Byung-Chul Kim
Electronics and Telecommunications Research Institute, Daejeon, KOREA
Chungnam National University, Daejeon, KOREA
hspark@etri.re.kr, jyl@cnu.ac.kr, byckim@cnu.ac.kr
- When HARQ BLER (Block Error Rate) is 10% and HARQ failure rate is 0.1% as LTE protocol design, if DL (Downlink) HARQ RTT is 8ms and TCP RTT is 10ms, TCP throughput is seriously decreased up to only 36% of maximum bandwidth. And if DL HARQ RTT is 16ms and TCP RTT is 10ms, TCP throughput is decreased up to only 19% of maximum bandwidth.
- To alleviate this problem, we propose "out-of-sequence delivery" in LTE link layer in order to decrease TCP RTT while HARQ or ARQ in LTE link layer is working for error recovery. The "out-of-sequence delivery" can decrease TCP RTT up to end-to-end RTT. While "out-of-sequence delivery" makes LTE link layer design simpler, but its throughput gain is considerable to the extent of 30% in average and 58% in maximum from our test results.
- If the RLC sub-layer receiver detects a gap in the sequence of the received PDUs, it starts a reordering timer assuming that the missing PDU still is being retransmitted in the HARQ protocol. HARQ failures appear if a maximum number of HARQ transmission attempts are exceeded or HARQ feedback NACK-to-ACK errors occur. When the timer expires, usually in a HARQ failure case, an RLC UM receiver delivers SDUs to PDCP with a certain amount of loss. However, an RLC AM receiver sends a status message comprising the sequence number of the missing PDUs to the sender. The ARQ function of the RLC AM sender performs retransmissions based on the received status message.
- The TCP RTT of packets which are contained PDUs from the gap SN to SN which received in-sequence, are proportional to the t_Reordering timer which is generally set as maximum HARQ transmission number times of MAC HARQ RTT
Let's put the mechanism in order. HARQ in MAC delivers transport blocks out of order, because each HARQ process retransmits on its own schedule. So the RLC receiver reorders the PDUs and starts t-Reordering at the first gap. Every PDU after the gap waits in the reception buffer, even when it arrived correctly. The TCP packets inside those PDUs therefore see the full wait as extra RTT.
3GPP later added an option for this in Release 15. TS 36.331 v19.3.0 carries rlc-OutOfOrderDelivery-r15 in RLC-Config-r15 and in RLC-Config-v1530. When it is configured, the RLC receiver in TS 36.322 delivers each RLC SDU to PDCP as soon as all its segments arrive, without waiting for the gap. The UE reports its support with rlc-AM-Ooo-Delivery and rlc-UM-Ooo-Delivery.
But this is not the same as the out-of-sequence delivery to TCP that the paper proposes. For a DRB with rlc-OutOfOrderDelivery, the PDCP t-Reordering-r12 field is mandatory, and TS 36.323 runs the PDCP reordering function for that bearer. So the in-sequence wait moves from RLC to PDCP, where its length is set by the PDCP timer.
In-sequence delivery turns a HARQ gap into TCP delay : every packet behind the gap waits until HARQ recovers it or t-Reordering expires.t-Reordering sets the length of the wait : it is usually sized to cover the HARQ retransmissions, so a longer HARQ RTT means a longer wait.Release 15 moved reordering but did not remove it : rlc-OutOfOrderDelivery-r15 skips RLC reordering, but the PDCP entity then reorders with its own t-Reordering.
Initial field performance measurements of LTE
How much does each antenna configuration add in a real channel? This 2008 Ericsson Review article measured layer-1 throughput for five antenna configurations, and it also compared UDP with FTP on the same link. The article is available as a pdf.
Jonas Karlsson, Mathias Riback
Ericsson
The two images below carry the article's Figure 6 to Figure 9. Figure 6 and Figure 7 plot layer-1 throughput against SNR at 10 MHz, first in the EVA 3 km/h channel model and then in a full-rank AWGN channel. Figure 8 compares UDP and FTP at 20 MHz with 2x2 in the Pedestrian B 3 km/h model. Figure 9 shows the CDF of throughput measured with dual-polarized antennas at 10 MHz.


More antennas raise the peak rate only at high SNR, and 4x4 MIMO gains most in a full-rank channel.
- Figure 6, EVA 3 km/h : at about 32 dB SNR, 4x4 MIMO reaches about 77 Mbps, 2x4 MIMO about 59 Mbps, 2x2 MIMO about 45 Mbps, 1x4 receive diversity about 35 Mbps and 1x2 receive diversity about 30 Mbps. Below about 10 dB, all five curves stay close together.
- Figure 7, full-rank AWGN : 4x4 MIMO reaches about 118 Mbps, while 2x4 and 2x2 MIMO both end near 73 to 75 Mbps. 1x4 and 1x2 receive diversity end near 38 to 39 Mbps.
- Figure 8, UDP and FTP : each modulation, QPSK, 16QAM and 64QAM at R=0.5, saturates at its own peak, about 27, 54 and 81 Mbps for UDP. FTP stays a little below UDP, and the gap is largest with 64QAM.
- Figure 9, CDF : 1x2 and 1x4 stay below about 40 Mbps. 2x2 and 2x4 spread up to about 75 to 80 Mbps, and 4x4 reaches about 100 Mbps.
Notice the difference between Figure 6 and Figure 7. The same 4x4 configuration gives about 77 Mbps in EVA and about 118 Mbps in full-rank AWGN. The fading channel does not always support four layers, so the extra antennas often add diversity instead of rank. Figure 9 shows the same effect in the field, where the 4x4 curve spreads over a wide range.
Figure 8 is useful for a throughput test on your own setup. UDP shows what the link delivers, while FTP runs over TCP and adds its own overhead and flow control. So when the FTP result is well below the UDP result on the same link, the TCP settings of the test are the first thing to check.
Antenna gain depends on channel rank : in full-rank AWGN, 4x4 MIMO reaches about 118 Mbps against about 73 Mbps for 2x2, but in EVA the gap is much smaller.Receive diversity does not raise the peak rate : in Figure 7, the 1x4 and 1x2 curves flatten at about the same single-layer peak.Compare UDP and TCP on the same link : a small gap is normal overhead, while a large gap points to the TCP setup.
Using iPerf to Troubleshoot Speed/Throughput Issues
A throughput test measures the link only when the traffic generator can fill the link. This 2011 blog post shows how much a single TCP stream with the default window can hide, and the same limit applies when the link under test is LTE. The post is available at this Link.
Posted by Andrew Tyler in Customer Service, SoftLayer, Technology, Tips and Tricks
- We were able to increase throughput from 29Mb/s with a single stream and the default TCP Window to 824Mb/s using a higher window and parallel streams. On a Gigabit link, this about the maximum throughput one could hope to achieve before saturating the link and causing packet loss.
- We will never get 100% out of any link. Typically, 90% utilization is about the real world maximum anyone will achieve. If you get any more, you'll begin to saturate the link and incur packet loss.
Let's see why the window matters so much. A TCP sender can have at most one window of unacknowledged data in flight. So a single stream cannot go faster than the window size divided by the RTT. For example, a 64 kB window with a 20 ms RTT limits one stream to about 26 Mbps, because 64 x 1024 x 8 bits divided by 0.02 s is about 26.2 Mbps.
An LTE link adds to that RTT. HARQ retransmissions, RLC reordering and the scheduling delay all appear as extra RTT, as the first section on this page shows. So on LTE, a larger window or several parallel streams is often needed before the test measures the radio link rather than the TCP window. iPerf sets the window with -w and the number of parallel streams with -P.
One TCP stream is limited to window divided by RTT : a small default window can hide most of the link capacity.Larger windows and parallel streams remove the TCP limit : the post went from 29 Mb/s to 824 Mb/s on a Gigabit link this way.Do not expect 100 percent utilization : about 90 percent is a practical maximum before packet loss starts.
Stochastic Forecasts Achieve High Throughput and Low Delay over Cellular Networks
Why do interactive applications work poorly on a cellular link that has plenty of average throughput? This talk starts from two measurements in live networks, and both concern variation on a short time scale rather than the average. The talk is available at this Link.
- Live Network Throughput is fluctuating in very wide range even in short time interval
- TCP turnaround time is fluctuating in very wide range even in short time interval
- Interactive Apps work poorly
- Possible solutions ?
The LTE side of this is easy to see. The eNB scheduler serves each UE at a rate that follows its channel quality and the load in the cell, and that rate can change from one TTI to the next. The eNB also keeps a buffer for each bearer. When the rate drops, the packets already in that buffer wait longer, so the RTT that TCP measures grows.
The title of the talk names the approach. The sender forecasts the link rate and limits what it sends, so that packets do not collect in the buffer. This is a change on the end hosts, and it needs no change in the LTE protocol stack.
Average throughput hides short-term variation : an interactive application feels the variation, not the average.Buffering turns rate variation into delay : packets queued at the eNB wait longer whenever the scheduled rate drops.The proposed fix sits in the end hosts : the sender adapts to a forecast of the link rate instead of filling the buffer.