عمليات الالتفاف
قالب صغير يمسح الصورة كاملة ليكشف الحواف والملامح والأشكال
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
Convolution slides a small matrix (the kernel) across an image, and at each position computes an element-wise product followed by a sum to produce a single output value. Repeating this “local weighted sum” at every location with shared weights lets it detect local patterns such as edges and textures using very few parameters.
الحدس المباشر
Picture a small stencil engraved with −1, 0 and +1 pressed against a patch of the photo: if the patch is a vertical edge, the weighted sum through the stencil comes out large; over flat sky it is near zero. Sweep the same stencil across the whole photo and you get a response map of “where edges of this kind are.” What the network learns is precisely the numbers on those stencils.
The geometry of convolution: a 5×5 patch through a 3×3 kernel (stride 1, no padding) yields a 3×3 response map in which the object boundary is extracted
Why images use convolution rather than fully connected layers: a comparison along locality, weight sharing and translation equivariance
- Convolution
- Fully connected
طريقة العمل
- 01
Slide and take the weighted sum
The kernel covers a small window of the input, multiplies element-wise and sums, producing one value on the output map. A larger kernel enlarges the receptive field but also raises parameters and compute.
- 02
Stride and padding control the output size
Stride sets how far the window jumps (downsampling); padding pads the border with zeros to preserve size. Together with the kernel size they fix the output dimensions: out = ⌊(in + 2·pad − kernel)/stride⌋ + 1.
- 03
Dilated and transposed convolution
Dilated convolution inserts gaps between kernel elements, enlarging the receptive field without new parameters and appearing widely in segmentation; transposed convolution goes the other way, upsampling feature maps to restore resolution or generate.
- 04
Stacking layers: from edges to objects
Shallow kernels learn edges and colour blobs, middle layers compose them into textures and parts, deep layers into objects. The key point: these kernels are not hand-designed — backpropagation learns them from data.
الصيغة الأساسية
out = ⌊(H + 2P − K) / S⌋ + 1The weights of an edge-detection kernel: a positive centre surrounded by negative neighbours, highly sensitive to abrupt brightness changes (hover a cell for its weight)
مجالات الاستخدام
- Feature-extraction backbones: the bottom of nearly every CNN, replacing hand-designed operators such as SIFT and HOG
- Edge and gradient detection: classic kernels like Sobel and Laplacian remain in use for image processing and industrial inspection
- Upsampling and generation: transposed convolution serves segmentation, super-resolution and decoders
- Large-receptive-field modelling: dilated convolution is used in semantic segmentation, audio and sequence modelling
مفاهيم خاطئة شائعة
- Convolution is not synonymous with “blurring”. Blur is merely one kernel among many; the same sliding mechanism with different weights sharpens, detects edges or responds to a specific orientation.
- Fewer parameters does not mean less computation. Weight sharing collapses the parameter count, but every location is still evaluated, so the cost grows with image area.
- Bigger kernels are not automatically better. Modern designs prefer stacking several 3×3 kernels to reach the same receptive field with fewer parameters and more nonlinearity.
مصطلحات أساسية
- Kernel / filter
- The small weight matrix that is learned
- Stride
- How many pixels the window jumps each step
- Padding
- Adding zeros at the border to control output size
- Receptive field
- The input region that one output pixel depends on