Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
1aurent29's comments
login
1aurent29
5 months ago
|
parent
|
context
|
next
[–]
| on:
MegaTrain: Full Precision Training of 100B+ Parame...
sounds very similar to
https://docs.pytorch.org/docs/stable/distributed.fsdp.fully_...
i wonder how much this could be replicated using only this pytorch primitive
RandyOrion
5 months ago
|
parent
|
next
[–]
Check out Fig. 6 in this paper, it shows the comparison between the proposed method and pytorch native FSDP offload method.
1aurent29
11 months ago
|
parent
|
context
|
prev
[–]
| on:
Which table format do LLMs understand best?
Common enough words like `function` and `class` are generally encoded as a single token by the tokenizer and may provide a slightly better context to the LLM. For openai you can test this stuff at
https://platform.openai.com/tokenizer
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: