Since the throughput process get involved in the full protocol stack and one/two PCs, and in addition IP tools, you would need to have tools to monitor each and every steps along the end-to-end data path. The logging tool should be able to show not only scheduling and event log, but also all the payload (contents of the data). Without these tools, you would end up saying "I have tested this device with many different test equipment and didn't see any problem before. This is the only equipment that I see this problem.. so the problem is on the equipment side." or "I have tested many different UE with this equipment, but I didn't see this kind of problem before. So this is UE side problem". Both may be right or wrong at the same time. Our job as an engineer is to find the root cause of the problem and fix it, not blaming the other party. (But to be honest, I have to admit I often blame the other party without even realizing it. Is this a kind of bad nature of engineers ? or my personal problem ?)
- Which tools do you need ?
- What do you check first in each log ?
- How do you narrow down the layer ?
- Reference
Which tools do you need ?
A throughput problem can sit in any layer between the two IP applications. So you need a view into each layer, and one tool rarely gives all of them. The list below gives one tool for each view, plus the two things that no tool gives you.
Let's try to have proper tools and skills to fight against the problem, not fight against your counter part engineers. My recommendation about the tool is as follows :
i) Ethernet, IP logging tool (e.g, Wireshark)
ii) UE side logging tool
iii) Network side logging tool
iv) YOUR SKILLs to use and analyze the logging tool
v) YOUR PATIENCE to step through the log for each and every transmission and reception
Each tool covers a different part of the data path listed on the Throughput Overview page. The Ethernet and IP logging tool covers the PC side, items i) to iii) and xii) to xiv). The UE side logging tool covers the UE protocol stack, items viii) to xi). The network side logging tool covers the test equipment or eNB stack, items iv) to vii).
Run the IP capture at both ends at the same time. Then you can compare what the server sent with what the UE side PC received, packet by packet. A packet that the server sent and the UE side never received was lost somewhere in between, and the radio side logs show where. A packet that the server never sent points to the server, the application or the PC network settings.
One tool per layer : the IP capture, the UE side log and the network side log each see a different part of the path.Capture both ends together : a packet-by-packet comparison tells you on which side the loss happens.Reading the log takes skill and time : items iv) and v) on the list are part of the tool set.
What do you check first in each log ?
A log is useful only if you know which fields to look at first. The table below gives a common starting point for each layer of an LTE data test, from the IP layer down to the PHY. It is a first check, not a complete list.
Layer and log | Check first | What it tells you |
IP, capture on both PCs | TCP retransmissions, duplicate ACKs, TCP window size | whether TCP limits the rate even when the radio link is clean |
IP, iperf UDP report | packet loss and jitter at a fixed offered rate | whether packets are lost below the IP layer |
RLC, UE or network log | STATUS PDUs with NACK_SN, expiry of t-PollRetransmit | whether RLC retransmits and where the gaps are |
MAC, UE or network log | HARQ ACK and NACK, Buffer Status Report, Power Headroom Report | HARQ failures, and whether the UE asks for enough UL grant and still has power |
PHY, UE or network log | reported CQI, RI and PMI, scheduled MCS and number of RBs, BLER | whether the scheduler actually uses the peak MCS, rank and RBs |
Read the table from the bottom up when the rate is low from the start. If the scheduled MCS, rank and RBs are already below the peak, the problem is in the radio link or the scheduler, and the upper layers are not the cause. If the PHY runs at the peak with a low BLER but the IP rate is low, move up the table to RLC, PDCP and TCP.
Start from the scheduled resources : MCS, rank and RBs set the ceiling for everything above them.RLC STATUS PDUs show the gaps : each NACK_SN is a missing RLC PDU that has to be retransmitted.BSR and PHR explain a low UL rate : a UE that reports little data or no power headroom gets small grants.
How do you narrow down the layer ?
Once you have the logs, the task is to find the layer where the rate drops. Let's compare the rate at each layer boundary, starting from the PHY and moving up. The rate should fall only by the header overhead at each step.
First, compute the PHY rate from the scheduled TBS in the log, for example with the method on the FDD Throughput Calculation page. Next, compare it with the MAC and RLC rates in the same log. Then compare those with the IP rate from the capture. A step larger than the header overhead marks the layer to look at.
Two quick tests help to split the problem further. A UDP test at a fixed rate removes TCP from the picture. So a good UDP result with a poor TCP result points to the TCP settings, the window size or the round trip time. Several parallel TCP streams give the same kind of hint. If they reach the rate that one stream cannot, one TCP stream is the limit, not the radio link.
Compare rates layer by layer : a drop larger than the header overhead shows the layer with the problem.Use UDP to remove TCP : a UDP result close to the peak clears the radio link and the lower layers.Parallel TCP streams expose a window limit : when more streams give more rate, a single stream is the bottleneck.
Reference
- 3GPP TS 36.321 v19.3.0 - Buffer Status Report and Power Headroom Report
- 3GPP TS 36.322 v19.0.0 - STATUS PDU, NACK_SN and t-PollRetransmit