Skip to main content
archive
Search Submit Donate Log in
Press Enter to search · Advanced search

arXiv
Trackbacks

Trackbacks indicate external web sites that link to articles in arXiv.org. Trackbacks do not reflect the opinion of arXiv.org and may not reflect the opinions of that article's authors.

Trackback guide

By sending a trackback, you can notify arXiv.org that you have created a web page that references a paper. Popular blogging software supports trackback: you can send us a trackback about this paper by giving your software the following trackback URL:

https://arxiv.org/trackback/{arXiv_id}

Some blogging software supports trackback autodiscovery -- in this case, your software will automatically send a trackback as soon as your create a link to our abstract page. See our trackback help page for more information.

Trackbacks for 2309.06180

How to co-design software/hardware architecture for AI/ML in a new era?

[ Towards Data Science - Medium@ INVALID-URL ] trackback posted Sun, 26 Nov 2023 15:49:13 UTC

Click to view metadata for 2309.06180

[Submitted on 12 Sep 2023]

Title:Efficient Memory Management for Large Language Model Serving with PagedAttention

Authors:Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica
Abstract:
Comments: SOSP 2023
Subjects: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as: arXiv:2309.06180 [cs.LG]
  (or arXiv:2309.06180v1 [cs.LG] for this version)
We gratefully acknowledge support from our major funders, member institutions, , and all contributors.
About · Help · Contact · Subscribe · Copyright · Privacy · Accessibility · Operational Status (opens in new tab)
Major funding support from
Simons Foundation Simons Foundation International Schmidt Sciences