Run this notebook online:\ |Binder| or Colab: |Colab|

.. |Binder| image:: https://mybinder.org/badge_logo.svg
   :target: https://mybinder.org/v2/gh/deepjavalibrary/d2l-java/master?filepath=chapter_attention-mechanisms/multihead-attention.ipynb
.. |Colab| image:: https://colab.research.google.com/assets/colab-badge.svg
   :target: https://colab.research.google.com/github/deepjavalibrary/d2l-java/blob/colab/chapter_attention-mechanisms/multihead-attention.ipynb

.. _sec_multihead-attention:

Multi-Head Attention
====================


In practice, given the same set of queries, keys, and values we may want
our model to combine knowledge from different behaviors of the same
attention mechanism, such as capturing dependencies of various ranges
(e.g., shorter-range vs. longer-range) within a sequence. Thus, it may
be beneficial to allow our attention mechanism to jointly use different
representation subspaces of queries, keys, and values.

To this end, instead of performing a single attention pooling, queries,
keys, and values can be transformed with :math:`h` independently learned
linear projections. Then these :math:`h` projected queries, keys, and
values are fed into attention pooling in parallel. In the end, :math:`h`
attention pooling outputs are concatenated and transformed with another
learned linear projection to produce the final output. This design is
called *multi-head attention*, where each of the :math:`h` attention
pooling outputs is a *head* :cite:`Vaswani.Shazeer.Parmar.ea.2017`.
Using fully-connected layers to perform learnable linear
transformations, :numref:`fig_multi-head-attention` describes
multi-head attention.

|Multi-head attention, where multiple heads are concatenated then
linearly transformed.| .. _fig_multi-head-attention:

Model
-----

Before providing the implementation of multi-head attention, let us
formalize this model mathematically. Given a query
:math:`\mathbf{q} \in \mathbb{R}^{d_q}`, a key
:math:`\mathbf{k} \in \mathbb{R}^{d_k}`, and a value
:math:`\mathbf{v} \in \mathbb{R}^{d_v}`, each attention head
:math:`\mathbf{h}_i` (:math:`i = 1, \ldots, h`) is computed as

.. math:: \mathbf{h}_i = f(\mathbf W_i^{(q)}\mathbf q, \mathbf W_i^{(k)}\mathbf k,\mathbf W_i^{(v)}\mathbf v) \in \mathbb R^{p_v},

where learnable parameters
:math:`\mathbf W_i^{(q)}\in\mathbb R^{p_q\times d_q}`,
:math:`\mathbf W_i^{(k)}\in\mathbb R^{p_k\times d_k}` and
:math:`\mathbf W_i^{(v)}\in\mathbb R^{p_v\times d_v}`, and :math:`f` is
attention pooling, such as additive attention and scaled dot-product
attention in :numref:`sec_attention-scoring-functions`. The multi-head
attention output is another linear transformation via learnable
parameters :math:`\mathbf W_o\in\mathbb R^{p_o\times h p_v}` of the
concatenation of :math:`h` heads:

.. math:: \mathbf W_o \begin{bmatrix}\mathbf h_1\\\vdots\\\mathbf h_h\end{bmatrix} \in \mathbb{R}^{p_o}.

Based on this design, each head may attend to different parts of the
input. More sophisticated functions than the simple weighted average can
be expressed.

.. |Multi-head attention, where multiple heads are concatenated then linearly transformed.| image:: https://raw.githubusercontent.com/d2l-ai/d2l-en/master/img/multi-head-attention.svg

.. code:: java

    %load ../utils/djl-imports
    %load ../utils/plot-utils
    %load ../utils/Functions.java
    %load ../utils/PlotUtils.java
    
    %load ../utils/attention/Chap10Utils.java
    %load ../utils/attention/DotProductAttention.java
    %load ../utils/attention/MultiHeadAttention.java
    %load ../utils/attention/PositionalEncoding.java

.. code:: java

    NDManager manager = NDManager.newBaseManager();

To allow for parallel computation of multiple heads, the below
``MultiHeadAttention`` class uses two transposition functions as defined
below. Specifically, the ``transposeOutput`` function reverses the
operation of the ``transposeQkv`` function.

.. code:: java

    public static NDArray transposeQkv(NDArray X, int numHeads) {
        // Shape of input `X`:
        // (`batchSize`, no. of queries or key-value pairs, `numHiddens`).
        // Shape of output `X`:
        // (`batchSize`, no. of queries or key-value pairs, `numHeads`,
        // `numHiddens` / `numHeads`)
        X = X.reshape(X.getShape().get(0), X.getShape().get(1), numHeads, -1);
    
        // Shape of output `X`:
        // (`batchSize`, `numHeads`, no. of queries or key-value pairs,
        // `numHiddens` / `numHeads`)
        X = X.transpose(0, 2, 1, 3);
    
        // Shape of `output`:
        // (`batchSize` * `numHeads`, no. of queries or key-value pairs,
        // `numHiddens` / `numHeads`)
        return X.reshape(-1, X.getShape().get(2), X.getShape().get(3));
    }
    
    public static NDArray transposeOutput(NDArray X, int numHeads) {
        X = X.reshape(-1, numHeads, X.getShape().get(1), X.getShape().get(2));
        X = X.transpose(0, 2, 1, 3);
        return X.reshape(X.getShape().get(0), X.getShape().get(1), -1);
    }

Implementation
--------------

In our implementation, we choose the scaled dot-product attention for
each head of the multi-head attention. To avoid significant growth of
computational cost and parameterization cost, we set
:math:`p_q = p_k = p_v = p_o / h`. Note that :math:`h` heads can be
computed in parallel if we set the number of outputs of linear
transformations for the query, key, and value to
:math:`p_q h = p_k h = p_v h = p_o`. In the following implementation,
:math:`p_o` is specified via the argument ``numHiddens``.

.. code:: java

    public static class MultiHeadAttention extends AbstractBlock {
    
        private int numHeads;
        public DotProductAttention attention;
        private Linear W_k;
        private Linear W_q;
        private Linear W_v;
        private Linear W_o;
    
        public MultiHeadAttention(int numHiddens, int numHeads, float dropout, boolean useBias) {
            this.numHeads = numHeads;
    
            attention = new DotProductAttention(dropout);
    
            W_q = Linear.builder().setUnits(numHiddens).optBias(useBias).build();
            addChildBlock("W_q", W_q);
    
            W_k = Linear.builder().setUnits(numHiddens).optBias(useBias).build();
            addChildBlock("W_k", W_k);
    
            W_v = Linear.builder().setUnits(numHiddens).optBias(useBias).build();
            addChildBlock("W_v", W_v);
    
            W_o = Linear.builder().setUnits(numHiddens).optBias(useBias).build();
            addChildBlock("W_o", W_o);
    
            Dropout dropout1 = Dropout.builder().optRate(dropout).build();
            addChildBlock("dropout", dropout1);
        }
    
        @Override
        protected NDList forwardInternal(
                ParameterStore ps,
                NDList inputs,
                boolean training,
                PairList<String, Object> params) {
            // Shape of `queries`, `keys`, or `values`:
            // (`batchSize`, no. of queries or key-value pairs, `numHiddens`)
            // Shape of `validLens`:
            // (`batchSize`,) or (`batchSize`, no. of queries)
            // After transposing, shape of output `queries`, `keys`, or `values`:
            // (`batchSize` * `numHeads`, no. of queries or key-value pairs,
            // `numHiddens` / `numHeads`)
            NDArray queries = inputs.get(0);
            NDArray keys = inputs.get(1);
            NDArray values = inputs.get(2);
            NDArray validLens = inputs.get(3);
            // On axis 0, copy the first item (scalar or vector) for
            // `numHeads` times, then copy the next item, and so on
            validLens = validLens.repeat(0, numHeads);
    
            queries =
                    transposeQkv(
                            W_q.forward(ps, new NDList(queries), training, params).get(0),
                            numHeads);
            keys =
                    transposeQkv(
                            W_k.forward(ps, new NDList(keys), training, params).get(0), numHeads);
            values =
                    transposeQkv(
                            W_v.forward(ps, new NDList(values), training, params).get(0), numHeads);
    
            // Shape of `output`: (`batchSize` * `numHeads`, no. of queries,
            // `numHiddens` / `numHeads`)
            NDArray output =
                    attention
                            .forward(
                                    ps,
                                    new NDList(queries, keys, values, validLens),
                                    training,
                                    params)
                            .get(0);
    
            // Shape of `outputConcat`:
            // (`batchSize`, no. of queries, `numHiddens`)
            NDArray outputConcat = transposeOutput(output, numHeads);
            return new NDList(W_o.forward(ps, new NDList(outputConcat), training, params).get(0));
        }
    
        @Override
        public Shape[] getOutputShapes(Shape[] inputShapes) {
            throw new UnsupportedOperationException("Not implemented");
        }
    
        @Override
        public void initializeChildBlocks(
                NDManager manager, DataType dataType, Shape... inputShapes) {
            try (NDManager sub = manager.newSubManager()) {
                NDArray queries = sub.zeros(inputShapes[0], dataType);
                NDArray keys = sub.zeros(inputShapes[1], dataType);
                NDArray values = sub.zeros(inputShapes[2], dataType);
                NDArray validLens = sub.zeros(inputShapes[3], dataType);
                validLens = validLens.repeat(0, numHeads);
    
                ParameterStore ps = new ParameterStore(sub, false);
    
                W_q.initialize(manager, dataType, queries.getShape());
                W_k.initialize(manager, dataType, keys.getShape());
                W_v.initialize(manager, dataType, values.getShape());
    
                queries =
                        transposeQkv(W_q.forward(ps, new NDList(queries), false).get(0), numHeads);
                keys = transposeQkv(W_k.forward(ps, new NDList(keys), false).get(0), numHeads);
                values = transposeQkv(W_v.forward(ps, new NDList(values), false).get(0), numHeads);
    
                NDList list = new NDList(queries, keys, values, validLens);
                attention.initialize(sub, dataType, list.getShapes());
                NDArray output = attention.forward(ps, list, false).head();
                NDArray outputConcat = Chap10Utils.transposeOutput(output, numHeads);
    
                W_o.initialize(manager, dataType, outputConcat.getShape());
            }
        }
    }


Let us test our implemented ``MultiHeadAttention`` class using a toy
example where keys and values are the same. As a result, the shape of
the multi-head attention output is (``batchSize``, ``numQueries``,
``numHiddens``).

.. code:: java

    int numHiddens = 100;
    int numHeads = 5;
    MultiHeadAttention attention = new MultiHeadAttention(numHiddens, numHeads, 0.5f, false);

.. code:: java

    int batchSize = 2;
    int numQueries = 4;
    int numKvpairs = 6;
    NDArray validLens = manager.create(new float[] {3, 2});
    NDArray X = manager.ones(new Shape(batchSize, numQueries, numHiddens));
    NDArray Y = manager.ones(new Shape(batchSize, numKvpairs, numHiddens));
    
    ParameterStore ps = new ParameterStore(manager, false);
    NDList input = new NDList(X, Y, Y, validLens);
    attention.initialize(manager, DataType.FLOAT32, input.getShapes());
    NDList result = attention.forward(ps, input, false);
    result.get(0).getShape();


.. parsed-literal::
    :class: output

    (2, 4, 100)


Summary
-------

-  Multi-head attention combines knowledge of the same attention pooling
   via different representation subspaces of queries, keys, and values.
-  To compute multiple heads of multi-head attention in parallel, proper
   tensor manipulation is needed.

Exercises
---------

1. Visualize attention weights of multiple heads in this experiment.
2. Suppose that we have a trained model based on multi-head attention
   and we want to prune least important attention heads to increase the
   prediction speed. How can we design experiments to measure the
   importance of an attention head?