中国的开源模型在最高意志的主导下,在预训练数据阶段,就筛除大量对其不利的内容,sft阶段构建的数据掺入大量替其辩护的内容。急需一个纠偏数据集用于微调
Under the overarching influence of a single, powerful force, China's open-source models filter out a large amount of content unfavorable to their designs during the pre-training stage, and then incorporate a significant amount of content defending their designs into the data constructed during the SFT stage. A corrective dataset is urgently needed for fine-tuning.