文章汇总
共 49 篇文章 · 约 55.8 万字 · 13 个分类 · 109 个标签 · 2025 年至今
2026 年 34 篇
- 代码语法测试
- 面向行为进化的持续强化学习研究
- Soft Actor-Critic 算法分析
- 拟实现方案
- 策略组合泛化思想
- Markdown 语法速查
- 本站是如何构建的
- 读书笔记:如何高效整理知识
- 自回归生成可以增强 dLLM 的探索性(The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models)
- 公式渲染:当数学遇上 Markdown
- 随机过程(11):马尔可夫链
- 随机过程(10):泊松过程(2): 过滤泊松过程
- 随机过程(9):泊松过程(1): 泊松分布
- 随机过程(8):高斯过程(4): 高斯过程的应用
- 随机过程(7):高斯过程(3): 高斯与非线性
- 实数集的完备性:七个等价表述及其证明
- 生成式推荐系统中的预测解码加速
- 优化器 (3)(Gram Newton-Schulz)
- 个性化搜索中的知识-动作对齐(KARMA: Knowledge-Action Regularized Multimodal Alignment for Personalized Search at Taobao)
- NVIDIA Nemotron 3 Super 技术解读(Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning)
- FlashAttention 系列
- 等周不等式:周长为定值时,面积最大的封闭图形是圆
- 随机过程(6):高斯过程(2): 多元高斯分布
- 变分序列级软策略优化(VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training)
- 随机过程(5):高斯过程(1): Gaussian is Everywhere
- 随机过程(4):多元相关
- 随机过程(3):非平稳随机过程
- 随机过程(2):宽平稳随机过程相关函数的时频分析
- 随机过程(1):线性相关
- RFT的熵动力学分析(On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models)
- 优化器 (2):Muon
- 优化器 (1):SGD和Adam
- 利用PRM监督生成式推荐模型(PROMISE: Process Reward Models Unlock Test-Time Scaling Laws in Generative Recommendations)
- 多目标优化对齐(HarmonRank: Ranking-aligned Multi-objective Ensemble for Live-streaming E-commerce Recommendation)
2025 年 15 篇
- 生成模型 (3.3):Flow Matching
- 生成模型 (3.2):Flow Model
- 生成模型 (3.1):Flow-based Method
- 生成模型 (2.1):Energy-based Model
- 生成模型 (1.3):Denoising Diffusion Probabilistic Model
- 生成模型 (1.2):Variational Auto-Encoder
- 生成模型 (1.1):变分推断
- 生成模型 (0):Overview of Deep Generative Modeling
- 通过拟合整个奖励分布进行强化学习(FlowRL: Matching Reward Distributions for LLM Reasoning)
- 强化学习基础 (3):动态规划求解
- 强化学习基础 (2):有限马尔可夫过程
- 强化学习中的熵 (2):熵安全策略
- 强化学习中的熵 (1):策略熵
- 统一精排阶段的特征交叉和序列建模(OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender)
- 强化学习基础 (1):多臂赌博机