RoPE: Properties, Patterns, and Long-Context Behavior
Rotary Position Embedding, usually called RoPE, is one of the most common positional encoding methods in modern decoder-only language mod
Rotary Position Embedding, usually called RoPE, is one of the most common positional encoding methods in modern decoder-only language mod
Stanford CS336 Assignment 1 is titled Building a Transformer LM . It covers the main ideas behind a decoder-only language model:
Discriminative Model: Learns the conditional probability <span class="
With self-supervised learning, we can train neural networks without the need for manually labelled datasets.
GPUs Modern AI training is structured as a hierarchy: individual GPUs (
在機器學習問題中,模型訓練的核心是梯度下降:
Natural Language Processing has undergone a massive evolution in recent years. To understand state-of-the-art models, we first need to look back at how we used to process sequences and the critical bottleneck that led to the invention of Attention.
Recurrent Neural Networks (RNNs) are a class of neural networks designed to handle sequential data. Unlike standa
Batch normalization and layer normalization improve training stability and reduce sensitivity to initialization by normalizing intermediate activations.
In this note, we will continue discussing CNNs.
When images get really large, we would have too many neurons to train in a regular neural network. The full connectivity is wasteful and the huge number of parameters would quickly lead to overfitting. What’s more, regular neural networks ignore the spatial structures of the image, missing out on some important information. As a result, a new type of neural network, the Convolutional Neural…
If you are diving into the mechanics of neural networks, you will inevitably encounter the backpropagation of the Softmax and Cross-Entropy loss. At first glance, the matrix calculus can feel a bit intimidating. However, once you break it down step-by-step using the chain rule, you will discover that the final gradients are incredibly elegant and intuitive.
The primary goal of backpropagation is to calculate the partial derivatives of the cost function C C C with respect to every weight W mathbf{W} W and bias b mathbf{b} b in the network.
Here we are introducing Neural Networks and Backpropagation.
With the score function and the loss function, now we focus on how we minimize the loss. Optimization is the process of finding the set of parameters W W W that minimize the loss function.
With the disadvantages of the KNN algorithm, we need to come up with a more powerful approach. The new approach will have two major components: a score function that maps the raw data to class scores, and a loss function that quantifies the agreement between the predicted scores and the ground truth labels.
Image Classification</h
In this part of the Cache Lab, the mission is simple yet devious: optimize matrix transposition for three specific sizes: 32x32, 64x64, and 61x67. Our primary enemy? Cache misses.
For the CSAPP Cache Lab, the students are asked to write a small C program (200~300 lines) that simulates a cache memory. The full
Here are the notes for CS188 Local Search.
Here are the lecture notes for UC Berkeley CS188 Lecture 3.
If you are a developer or power user on a Mac, you probably type your password into the terminal dozens of times a day using sudo . If your Mac has a TouchID sensor, you can save time and keystrokes by configuring your terminal to accept your fingerprint instead of your password. Here is a quick guide on how to set it up.
Here are the lecture notes for UC Berkeley CS188 Lecture 1 and 2.
前幾天才猛然驚覺,按照慣例,是時候寫年終總結了。盯著白花花的螢幕,半天拼湊不出一句話。似乎 2025 是注定被遺忘的一年。 既然這是難以回憶難以定義的一年,不如我們留白。 <a c
做完了 CSAPP Bomb Lab,寫一篇解析。 題目要求 運行一個二進制文件
本文匯總 x64 架構下最核心的暫存器狀態與 ABI 約定。 暫存器 x64 暫存器
前一段時間做完了 CSAPP 的第一個 Lab,寫一篇總結。(其實這篇文章拖了很久) CS:APP Data Lab 旨在通過一系列位操作謎題,訓練對整數和浮點數底層表示(特別是補碼和 IEEE 754 標準)的
矩陣的 QR 分解在電腦運算中可能造成誤差,本文探討一下一種改進版本的 Gram-Schmidt 正交化方法。 經典的 Gram-Schmidt 方法可能造成數值不穩定性。在電腦中,舍入誤差可能會累積,造成得到的
掩碼 (Mask) 是一種位運算技巧,它使用一個特定的值(掩碼)與目標值進行 <math xmlns="h
在做 CSAPP Data Lab 的時候,關於整數溢位,遇到一些問題。 題幹 <figure
最近在複習資料結構與演算法,聊聊快速排序的幾種劃分演算法。 快速排序思路 <p
此生此夜不長好,明月明年何處看? 雨巷獨行,秉燭對月。我走遍了整座城市,卻再也找不到那個獨特的角落。或許只是那盞燈或那半掩的門扉,在人生之中卻像風暴裡海上的信標,似這江南春雨中雲霧朦朧處,簷下的紙燈。舊夢重溫,幼
憶起一場夢 太累了 寫得很爛 使用 DeepSeek 處理文章中的一些重複的段落 青石階蜿蜒向上,盡頭處懸著座褪色的廟宇。香火薰染的簷角垂著銅鈴,風起時聲響像隔世的嘆息。你隨人潮挪動腳步,鞋底碾碎階縫裡冒出的野蕨
你乘火車回到故鄉。 車廂裡昏暗的燈光,顫抖著拼寫出別離的感傷。窗外是多少山川歲月。傾倒的電線杆,望不到盡頭的平原,水壩,大江東去,沿著車門一直伸展向世界的盡頭。你所知道的只有這窗外的世界。多少片雪花落下,你不知道
紀念一段往事 <a class="markdownIt-Anchor" href="#橋樑-小滿-202421-
舊詩一首 淺灘上 一群猴子說教 太陽未生的角落 玻璃杯破碎 火球燃燒 門外守著 枯萎的向日葵 熟稔的幽谷<br
獻給 2024,和夢中的水鄉。 2024年的盡頭,希望寫下這些文字,腦海中卻是虛空,難以回想起2024年發生的一切。這一年,太多變化,太多茫然,太多草率。選擇
太陽 請給冬日動人的悲憫 你是空有光明的夜燈 無力帶給寒天一絲暖意 太陽 請把生命留在白晝 你是將遭審判的囚徒 決絕將赴星月的刑場
浪宣告朗,日光下, 我試著抓住大海 —— 它 卻後退,化身成點點泡沫。 2023.8.10 《泡沫》 從夜夢中驚醒,已是黃
題目連結 <a class="markdownIt-Anchor" href="#題面"
題目連結: P5888 <a class="markdownIt-Anchor" href
“我的生命將是大海。” 青島車站,至今仍然保留著殖民地時期的風貌。磚紅色的屋頂,靜立在月色海風中,沉醉在浪濤的朦朧樂聲中——那是不遠處棧橋邊的大海。候車大廳靜悄悄,耳畔不時傳來廣播的聲響,也似輕煙一
懶得寫簡介,你們自己點進去看 doge 初秋的幽風再不能喚醒, 心中將乾涸的冰流, 直到那西湖——野徑清夜的迴風 舞雪在我心中翩躚。 煙霞嶺的桂香,
《飢餓表演藝術家》是卡夫卡曾計劃印刷的短篇故事合集,他死後才出版。其中與合集同名的小說,也是最受他珍重的幾個短篇小說之一。 故事情節大致如下: 多年前流行著一種飢餓藝術,人們爭著來觀看。飢餓藝術家往
最近要用 Python 做一些小專案,記錄一些學習心得。以及這個部落格再不更新技術文章,就變成文學部落格了,顯然和初衷相違背() <a class="markdownIt-Anchor" hr
只恨一切都成了泡沫…… 浪宣告朗,日光下, 我試著抓住大海——它 卻後退,化身成點點泡沫。 海波泛起星光閃爍,<br /
“救救孩子……” 漸弱的呼喊,魯迅在狂人的絕望中收束全篇。這種無力卻格外讓人感覺似有一股刺骨之寒,或能於鐵屋中驚起更多較清醒的“不幸者”。 <a class="markdownIt-Anc
好久不見啊! 古羅馬的鬥獸場,數千年矗立在地中海沿岸的小城。濛濛細雨,如遊絲,洗濯時間長河中穿梭著的風塵僕僕的旅客。 或許想起在煙雨江南,於靉靆之下縵立,遠眺城牆之外。人們總是刻意忽視時間,以為出了
秋潮漲落, 在那心的橋樑。 彼岸還有草露甘香。 春泥, 是明天的味道。 明天?回答今日的困惑—— 海的女兒, 帶來獨木扁舟, 和
一 六月的鳴蟬嘒嘒,微弱卻唱響夜空。 路旁的小樓上,沐心倚著陽臺的欄杆,遠眺萬里澄澈的夜空。