x = np.arange(12)
xDive into Deep Learning · §1.1
Storing & transforming data with tensors
The n-dimensional arrays that every model in this book is built on.
Motivation
ndarray.Rank = number of axes; shape = size per axis.
01
Getting Started
creating & inspecting tensors
Getting Started
Getting Started
For weight init, randn draws from \mathcal{N}(0, 1):
array([[-1.0732179 , 1.4494689 , -0.9780016 , -0.69197655],
[ 0.30483028, 2.2523196 , -0.2745827 , 0.2653999 ],
[ 2.984504 , 0.7354883 , -1.3070105 , 0.9972589 ]])
Also zeros, ones, full(shape, value), eye(n). Random values break symmetry when initializing network weights; lists let you type a tensor by hand.
Getting Started
02
Indexing & Slicing
reading & writing elements, rows, ranges
Indexing & Slicing
Indexing & Slicing
03
Operations
elementwise math, joins, comparisons, broadcasting
Operations
The operators + - * / ** act elementwise on matching shapes:
(array([ 3., 4., 6., 10.]),
array([-1., 0., 2., 6.]),
array([ 2., 4., 8., 16.]),
array([0.5, 1. , 2. , 4. ]),
array([ 1., 4., 16., 64.]))
Any scalar→scalar map (exp, sin, log) extends to a whole tensor.
Operations
Operations
Comparisons return a boolean tensor.
A ready-made mask:
array([[False, True, False, True],
[False, False, False, False],
[False, False, False, False]])
==, <, > build masks; sum, mean, max collapse axes; add dim= to reduce just one.
Operations · the exception
Size-1 axes are virtually stretched
a 3\times1 plus a 1\times2 gives a 3\times2:
array([[0., 1.],
[1., 2.],
[2., 3.]])
Any axis of size 1 stretches to match the other tensor, without a copy.
Compatible only if each axis is equal or 1.
Operations · the exception
Line up (3, 2) and (2, 3) from the right, pairing 2 with 3 and 3 with 2: no pair matches, neither member is 1, so the framework raises rather than guessing:
Traceback (most recent call last):
File "/home/smola/mxnet/src/operator/numpy/./../tensor/elemwise_binary_broadcast_op.h", line 69
MXNetError: Check failed: l == 1 || r == 1: operands could not be broadcast together with shapes [3,2] [2,3]
Broadcasting aligns shapes from the right; each axis pair must be equal or 1.
04
Memory & Interop
in-place updates and leaving the tensor world
Y = Y + XPerformance
Every arithmetic expression allocates a new tensor
costly when Y is gigabytes and updated many times per second:
False
id(Y) changed: Y is now bound to a new tensor object.
Performance
Interop
Convert to / from a NumPy ndarray:
(numpy.ndarray, mxnet.numpy.ndarray)
The result is a copy; host/device arrays don’t share storage here.
Wrap-up
arange, zeros, ones, randn, tensor([…])..shape, .numel(), reshape.cat.X[:] = …, +=), or in JAX via jit buffer reuse..item() for scalars.