Skip to content
TH-20 · MSc Photonics thesis · 2020

MSc Photonics · 2020

Machine learning enabled nonlinear phase noise mitigation for coherent optical systems

Glass bends bright light. A receiver that learns its decisions from examples reads the twisted signal anyway, and the link reaches further.

Lund University, Atomic Physics Division. Research done at KTH and the RISE Kista High-speed Transmission Lab, Stockholm. Supervised by Cord Arnold (Lund), Sergei Popov (KTH).

Cite: S. S. Gutiérrez Pulgarín, Machine learning enabled nonlinear phase noise mitigation for coherent optical systems, MSc Photonics thesis, Lund University, 2020.

FIG. 1 ·The thesis in one row: 16-QAM at 1 mW after 0, 16, 32, 48 and 63 spans of 80 km. Each point is a received symbol, coloured by the power of the symbol sent: blue the four inner, teal the eight middle, green the four outer. The brighter the symbol, the further the Kerr effect turns it. Drawn from the thesis simulation files

01 The question

Long fibre links carry data as light whose amplitude and phase both encode bits. Push more power into the fibre and the signal gets cleaner, until the glass itself starts to bend the light: the brighter a symbol, the more its phase turns. Amplifier noise makes that turn random. This nonlinear phase noise is the wall that limits how far a link can reach.

The usual fix needs a model of the fibre. The thesis asks whether a receiver can learn its decisions from the received symbols alone, with support vector machines and random forests, and no knowledge of the link.

02 Light as a carrier

A coherent transmitter writes each symbol as one point in the plane of amplitude and phase, the constellation. More points carry more bits per symbol, but they sit closer together, so less noise is enough to push a symbol across into its neighbour's territory.

FIG. 2 ·Modulation formats, from 1 bit per symbol to 4. Thesis Fig. 1.3.2

Ordinary noise blurs every point the same way, in every direction. A receiver handles it with straight decision lines halfway between the points.

FIG. 3 ·16-QAM under linear noise, simulated. Thesis Fig. 1.4.1

03 When the glass bends the light

At high power the fibre's refractive index depends on the light's own intensity: the Kerr effect. Each symbol picks up a phase turn proportional to its power, so the outer points of a constellation turn further than the inner ones. Noise added by every amplifier changes the power a little, and the Kerr effect turns that into random phase. This is the Gordon–Mollenauer effect, and the constellation becomes a spiral that straight lines cannot separate.

16-QAM twisted into a spiral, outer clusters turned further than inner ones (full size, opens in a new tab)
(a) Gordon–Mollenauer effect
16-QAM with clusters spread by neighbouring channels (full size, opens in a new tab)
(b) Cross-phase modulation
FIG. 4 ·16-QAM under nonlinear noise, simulated. Thesis Fig. 1.4.2

Try it below. Pick a format, add distance, raise the power. Too little power and amplifier noise wins; too much and the spiral wins. Then switch the receiver to learned regions.

Format
Receiver
Bits per symbol
4
Wrong, nearest
291 / 1024
Wrong, learned
14 / 1024
Spiral
1.03 rad

+ sent● received● decided wrongShading: decision regions

FIG. 5 ·Nonlinear phase noise, interactive. Illustrative model in normalised units, not the thesis's simulation; the learned receiver is a k-nearest-neighbour stand-in for the thesis's SVM and random forest

Three things to notice in the model. QPSK barely suffers: all its points share one ring, so they turn together and the receiver takes the turn out. 64-QAM carries three times the bits and breaks first. 16-APSK puts its 16 points on two rings, which holds up better than the square grid under the same noise.

04 Learning the decision

Instead of modelling the fibre, the thesis trains a classifier on received symbols whose values are known, then lets it draw the decision regions. The regions bend to follow the spiral. Two classifiers learn the same data differently: the random forest cuts the plane into boxes, the support vector machine draws smooth curves.

Coloured decision regions made of rectangles, following a twisted constellation (full size, opens in a new tab)
(a) Random forest
Coloured decision regions with smooth curved borders, following a twisted constellation (full size, opens in a new tab)
(b) Support vector machine
FIG. 6 ·Decision regions learned from 16-QAM after 26 spans of 80 km. Thesis Fig. 1.5.1
FIG. 7 ·16-QAM after 63 spans at 1 mW, simulated: the received symbols and the decision regions learned from them. From the thesis code repository

05 The simulation

A single-channel 16-QAM link at 128 Gbit/s, simulated in VPI: a transmitter with adjustable launch power, a loop of 80 km dispersion-shifted fibre spans with low-noise amplifiers, and a coherent receiver. The loop runs up to 64 times, 5120 km. Amplifier noise is kept small so the nonlinear effects dominate. A link counts as reliable while its symbol error rate stays below 10⁻⁵.

FIG. 8 ·The simulated link: transmitter, fibre loop, coherent receiver. Thesis Fig. 4.2.1
16-QAM clusters stretched into arcs, several of them crossing the straight-edged decision cells drawn around the cluster centres (full size, opens in a new tab)
(a) With nonlinear phase noise
The same 16-QAM after compensation: round clusters, almost all inside their own decision cells (full size, opens in a new tab)
(b) After compensation
FIG. 9 ·Square 16-QAM with a mean nonlinear phase shift of 0.13 mrad, before and after electronic compensation with optimal scaling factor α (scheme of K.-P. Ho and J. Kahn, thesis ref. [7]). Thesis Fig. 3.1.1

Each launch power gives three error curves against distance: the raw signal, a compensation scheme derived from the physics of the link, and the learned receiver. The vertical lines mark where each one stops being reliable.

Launch power
110⁻³10⁻⁶10⁻⁹10⁻¹²
10⁻⁵ limitOn the axis: below 10⁻¹²
010002000300040005000

Y: symbol error rate · X: distance, km● Raw■ Compensated┆ Reach

Reach, raw
160 km
Reach, comp.
400 km
Ratio
2.5×
Show the numbers
PowerCompensated reachRaw reachRatio
0.17 mW1600 km1040 km1.5×
0.25 mW1360 km720 km1.9×
0.5 mW800 km400 km2.0×
1 mW400 km160 km2.5×
1.5 mW560 km400 km1.4×
2 mW560 km400 km1.4×
FIG. 10 ·Symbol error rate against distance at six launch powers, from the simulation files: raw signal and analytical compensation. Reach is read where the rate first reaches 10⁻⁵. The learned receiver is in the next figure, thesis Fig. 4.2.4. Distances are 64 spans of 80 km, to 5120 km (the files label them in 88 km steps, a slip in the simulation script). With 4096 symbols per frame, rates below about 2 × 10⁻⁴ are the simulator's estimates, not counted errors
FIG. 11 ·Error rate against distance at 1 mW: raw signal, analytical compensation and learned receiver. Orange: raw signal; blue: analytical compensation; green: learned receiver. Thesis Fig. 4.2.4

06 Key figures

ParameterValueUnit
Format, single channel16-QAM, 128Gbit/s
Fibre span80km
Longest link simulated, 64 spans5120km
Reliable linkSER < 10⁻⁵
Launch power, learned-receiver comparison (thesis)1mW
Reach gain, learned receiver, simulated2.5 simulated; up to 9.4 by exponential fit×
WDM test system, 5 wavelengths × 2 polarisations (10 channels)1280Gbit/s
Channel spacing, WDM50GHz

07 Result

On the long single-channel link, the learned receiver matches the compensation scheme that knows the link, with no information about the link at all, and stretches reach 2.5 times. An exponential fit to its error rate suggests up to 9.4 times; the simulation ran 4096 symbols per frame, too few to count rarer errors directly. On the ten-channel WDM system no clear gain shows: noise from neighbouring channels changes in time, and a fixed decision map cannot follow it.

The conclusion: on a long-haul link, knowledge of the system can be replaced by learning.

08 Where it leads now

The idea I still like from this thesis: you do not have to model the Kerr effect, or know every property of the fibre, to undo what it does. Send a known pattern first, let the receiver learn the distortion, and each span calibrates itself quickly. The same idea, learning the distortion instead of modelling it, took me into radio, where I worked on digital pre-distortion for power amplifiers.