dgwon commited on
Commit
d300f4e
·
verified ·
1 Parent(s): 2223547

Upload folder using huggingface_hub

Browse files
1_Pooling/config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "word_embedding_dimension": 384,
3
+ "pooling_mode_cls_token": false,
4
+ "pooling_mode_mean_tokens": true,
5
+ "pooling_mode_max_tokens": false,
6
+ "pooling_mode_mean_sqrt_len_tokens": false,
7
+ "pooling_mode_weightedmean_tokens": false,
8
+ "pooling_mode_lasttoken": false,
9
+ "include_prompt": true
10
+ }
README.md ADDED
@@ -0,0 +1,710 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ tags:
3
+ - sentence-transformers
4
+ - sentence-similarity
5
+ - feature-extraction
6
+ - dense
7
+ - generated_from_trainer
8
+ - dataset_size:26136
9
+ - loss:CachedMultipleNegativesRankingLoss
10
+ base_model: intfloat/e5-small-v2
11
+ widget:
12
+ - source_sentence: 'query: Do Artificial Food Dyes Cause ADHD?'
13
+ sentences:
14
+ - 'passsage: It has become widely accepted that Type 2 diabetes is inevitably life-long,
15
+ with irreversible and progressive beta cell damage. However, the restoration of
16
+ normal glucose metabolism within days after bariatric surgery in the majority
17
+ of people with Type 2 diabetes disproves this concept. There is now no doubt that
18
+ this reversal of diabetes depends upon the sudden and profound decrease in food
19
+ intake, and does not relate to any direct surgical effect. The Counterpoint study
20
+ demonstrated that normal glucose levels and normal beta cell function could be
21
+ restored by a very low calorie diet alone. Novel magnetic resonance methods were
22
+ applied to measure intra-organ fat. The results showed two different time courses:
23
+ a) resolution of hepatic insulin sensitivity within days along with a rapid fall
24
+ in liver fat and normalisation of fasting glucose levels; and b) return of normal
25
+ beta cell insulin secretion over weeks in step with a fall in pancreas fat. Now
26
+ that it has been possible to observe the pathophysiological events during reversal
27
+ of Type 2 diabetes, the reverse time course of events which determine the onset
28
+ of the condition can be identified. The twin cycle hypothesis postulates that
29
+ chronic calorie excess leads to accumulation of liver fat with eventual spill
30
+ over into the pancreas. These self-reinforcing cycles between liver and pancreas
31
+ eventually cause metabolic inhibition of insulin secretion after meals and onset
32
+ of hyperglycaemia. It is now clear that Type 2 diabetes is a reversible condition
33
+ of intra-organ fat excess to which some people are more susceptible than others.'
34
+ - 'passsage: A great deal of effort is now being devoted to the development of new
35
+ drugs that hopefully will control the spread of inoperable cancer by safely inhibiting
36
+ tumor-evoked angiogenesis. However, there is growing evidence that certain practical
37
+ nutritional measures have the potential to slow tumor angiogenesis, and it is
38
+ reasonable to anticipate that, by combining several measures that work in distinct
39
+ but complementary ways to impede the angiogenic process, a clinically useful ''multifocal
40
+ angiostatic therapy'' (MAT) might be devised. Several measures which might reasonably
41
+ be included in such a protocol are discussed below, and include: a low-fat, low-glycemic
42
+ index vegan diet, which may down-regulate the systemic IGF-I activity that supports
43
+ angiogenesis; supplemental omega-3-rich fish oil, which has been shown to inhibit
44
+ endothelial expression of Flk-1, a functionally crucial receptor for VEGF, and
45
+ also can suppress tumor production of pro-angiogenic eicosanoids; high-dose selenium,
46
+ which has recently been shown to inhibit tumor production of VEGF; green tea polyphenols,
47
+ which can suppress endothelial responsiveness to both VEGF and fibroblast growth
48
+ factor; and high-dose glycine, whose recently reported angiostatic activity may
49
+ reflect inhibition of endothelial cell mitosis, possibly mediated by activation
50
+ of glycine-gated chloride channels. In light of evidence that tumor-evoked angiogenesis
51
+ has a high requirement for copper, copper depletion may have exceptional potential
52
+ as an angiostatic measure, and is most efficiently achieved with the copper-chelating
53
+ drug tetrathiomolybdate. If logistical difficulties make it difficult to acquire
54
+ this experimental drug, high-dose zinc supplementation can achieve a slower depletion
55
+ of the body''s copper pool, and in any case can be used as maintenance therapy
56
+ to maintain an adequate level of copper depletion. A provisional protocol is offered
57
+ for a nutritionally based MAT entailing a vegan diet and supplemental intakes
58
+ of fish oil, selenium, green tea polyphenols, glycine, and zinc. Inasmuch as cox-2
59
+ is overexpressed in many cancers, and cAMP can boost tumor production of various
60
+ angiogenic factors as well as autogenous growth factors, adjunctive use of cox-2-specific
61
+ NSAIDS may be warranted in some cases.'
62
+ - 'passsage: Attention deficit hyperactivity disorder (ADHD) is one of the most
63
+ common behavioral disorders in children. Symptoms of ADHD include hyperactivity,
64
+ low frustration tolerance, impulsivity, and inattention. While the biological
65
+ pathways leading to ADHD are not clearly delineated, a number of genetic and environmental
66
+ risk factors for the disorder are recognized. In the early 1970s, research conducted
67
+ by Dr. Benjamin Feingold found that when hyperactive children were given a diet
68
+ free of artificial food additives and dyes, symptoms of hyperactivity were reduced.
69
+ While some clinical studies supported these findings, more rigorous empirical
70
+ studies conducted over the next 20 years were less positive. As a result, research
71
+ on the role of food additives in contributing to ADHD waned. In recent years,
72
+ however, interest in this area has revived. In response to more recent research
73
+ and public petitions, in December 2009 the British government requested that food
74
+ manufacturers remove most artificial food dyes from their products. While these
75
+ strictures could have positive effects on behavior, the removal of food dyes is
76
+ not a panacea for ADHD, which is a multifaceted disorder with both biological
77
+ and environmental underpinnings. © 2011 International Life Sciences Institute.'
78
+ - source_sentence: 'query: Is HDL High in Active Africans?'
79
+ sentences:
80
+ - 'passsage: The serum concentration of high-density lipoprotein cholesterol and
81
+ the proportion it constitutes of total serum cholesterol are high in children
82
+ and low in sufferers from coronary heart disease (CHD). Studies in elderly black
83
+ Africans in Western Transvaal showed them to be free of CHD. HDL concentrations
84
+ measured at birth and in groups of 10- to 12-year-olds, 16- to 18-year olds, and
85
+ 60- to 69-year-olds showed mean values of 0.96, 1.71, 1.58, and 1.94 mmol/l (36,
86
+ 66, 61, and 65 mg/100 ml) respectively; these concentrations constitued about
87
+ 56%, 54%, and 45%, and 47%, of total cholesterol. Values thus did not fall from
88
+ youth to age as they did in whites. Rural South African blacks live on a diet
89
+ high in fibre and low in animal protein and fat; children are active; and adults
90
+ remain active even when old. These high values of HDL may well be representative
91
+ for a population that is active, used to a frugal traditional diet, and free from
92
+ CHD.'
93
+ - 'passsage: Flavonoids have been hypothesized to reduce cancer risk. Previous epidemiological
94
+ studies conducted to evaluate this hypothesis have not assessed all flavonoids,
95
+ including classes that could contribute to intake among Americans, which would
96
+ result in an underestimation of intake. This misclassification could mask variability
97
+ among individuals, resulting in attenuated effect estimates for the association
98
+ between flavonoids and cancer. To augment flavonoid and lignan intake estimates,
99
+ we developed a database that can be used in conjunction with a food-frequency
100
+ questionnaire (FFQ). Coupling information derived from the available literature
101
+ with the U.S. Department of Agriculture databases, we estimated content of 6 flavonoid
102
+ classes and lignans for 50 food group items. We combined these estimates with
103
+ responses from a modified Block FFQ that was self-completed in 1996-1997 by a
104
+ population-based sample of women without breast cancer on Long Island, New York
105
+ (n = 1,500). Total flavonoid and lignan content of food items ranged from 0 to
106
+ 129 mg/100 g, and the richest sources were tea, cherries, and grapefruit. Individual
107
+ intake estimates, from highest to lowest, were flavan-3-ols, flavanones, flavonols,
108
+ lignans, isoflavones, anthocyanidins, and flavones. Each class of flavonoids and
109
+ lignans exhibited a wide range of intake levels. This database is useful to quantify
110
+ flavonoid and lignan intake for other observational studies conducted in the United
111
+ States that utilize the Block FFQ.'
112
+ - 'passsage: Several prospective studies considered the relation between coffee
113
+ consumption and mortality. Most studies, however, were underpowered to detect
114
+ an association, since they included relatively few deaths. To obtain quantitative
115
+ overall estimates, we combined all published data from prospective studies on
116
+ the relation of coffee with mortality for all causes, all cancers, cardiovascular
117
+ disease (CVD), coronary/ischemic heart disease (CHD/IHD) and stroke. A bibliography
118
+ search, updated to January 2013, was carried out in PubMed and Embase to identify
119
+ prospective observational studies providing quantitative estimates on mortality
120
+ from all causes, cancer, CVD, CHD/IHD or stroke in relation to coffee consumption.
121
+ A systematic review and meta-analysis was conducted to estimate overall relative
122
+ risks (RR) and 95 % confidence intervals (CI) using random-effects models. The
123
+ pooled RRs of all cause mortality for the study-specific highest versus low (≤1
124
+ cup/day) coffee drinking categories were 0.88 (95 % CI 0.84-0.93) based on all
125
+ the 23 studies, and 0.87 (95 % CI 0.82-0.93) for the 19 smoking adjusting studies.
126
+ The combined RRs for CVD mortality were 0.89 (95 % CI 0.77-1.02, 17 smoking adjusting
127
+ studies) for the highest versus low drinking and 0.98 (95 % CI 0.95-1.00, 16 studies)
128
+ for the increment of 1 cup/day. Compared with low drinking, the RRs for the highest
129
+ consumption of coffee were 0.95 (95 % CI 0.78-1.15, 12 smoking adjusting studies)
130
+ for CHD/IHD, 0.95 (95 % CI 0.70-1.29, 6 studies) for stroke, and 1.03 (95 % CI
131
+ 0.97-1.10, 10 studies) for all cancers. This meta-analysis provides quantitative
132
+ evidence that coffee intake is inversely related to all cause and, probably, CVD
133
+ mortality.'
134
+ - source_sentence: 'query: Do Flavonoids Help Reduce Heart Disease?'
135
+ sentences:
136
+ - 'passsage: Tea (Camellia sinensis, Theaceae) and tea polyphenols have been studied
137
+ for the prevention of chronic diseases, including obesity. Obesity currently affects
138
+ >20% of adults in the United States and is a risk factor for chronic diseases
139
+ such as type II diabetes, cardiovascular disease, and cancer. Given this increasing
140
+ public health concern, the use of dietary agents for the prevention of obesity
141
+ would be of tremendous benefit. Whereas many laboratory studies have demonstrated
142
+ the potential efficacy of green or black tea for the prevention of obesity, the
143
+ underlying mechanisms remain unclear. The results of human intervention studies
144
+ are mixed and the role of caffeine has not been clearly established. Finally,
145
+ there is emerging evidence that high doses of tea polyphenols may have adverse
146
+ side effects. Given that the results of scientific studies on dietary components,
147
+ including tea polyphenols, are often translated into dietary supplements, understanding
148
+ the potential toxicities of the tea polyphenols is critical to understanding their
149
+ potential usefulness in preventing obesity. In this review, we will critically
150
+ evaluate the evidence for the prevention of obesity by tea, discuss the relevance
151
+ of proposed mechanisms in light of tea polyphenol bioavailability, and review
152
+ the reports concerning the toxic effects of high doses of tea polyphenols and
153
+ the implication that this has for the potential use of tea for the prevention
154
+ of obesity. We hope that this review will expose areas for further study and encourage
155
+ research on this important public health issue.'
156
+ - 'passsage: Background & Aims The historical prevalence and long-term outcome of
157
+ undiagnosed celiac disease (CD) are unknown. We investigated the long-term outcome
158
+ of undiagnosed CD and whether the prevalence of undiagnosed CD has changed during
159
+ the past 50 years. Methods This study included 9,133 healthy young adults at Warren
160
+ Air Force Base (sera were collected between 1948 and 1954) and 12,768 sex-matched
161
+ subjects from 2 recent cohorts from Olmsted County, Minnesota, with either similar
162
+ years of birth (n=5,558) or age at sampling (n=7,210) to that of the Air Force
163
+ cohort. Sera were tested for tissue transglutaminase and, if abnormal, for endomysial
164
+ antibodies. Survival was measured during a follow-up period of 45 years in the
165
+ Air Force cohort. The prevalence of undiagnosed CD between the Air Force cohort
166
+ and recent cohorts was compared. Results Of 9,133 persons from the Air Force cohort,
167
+ 14 (0.2%) had undiagnosed CD. In this cohort, during 45 years of follow-up, all-cause
168
+ mortality was greater in persons with undiagnosed CD than among those who were
169
+ seronegative (hazard ratio=3.9; 95% CI, 2.0–7.5; P<.001). Undiagnosed CD was found
170
+ in 68 (0.9%) persons with similar age at sampling and 46 (0.8%) persons with similar
171
+ years of birth. The rate of undiagnosed CD was 4.5-fold and 4-fold greater in
172
+ the recent cohorts (respectively) than in the Air Force cohort (both P ≤ .0001).
173
+ Conclusions During 45 years of follow-up, undiagnosed CD was associated with a
174
+ nearly 4-fold increased risk of death. The prevalence of undiagnosed CD appears
175
+ to have increased dramatically in the United States during the past 50 years.'
176
+ - 'passsage: Dietary flavonols and flavones are subgroups of flavonoids that have
177
+ been suggested to decrease the risk of coronary heart disease (CHD). The authors
178
+ prospectively evaluated intakes of flavonols and flavones in relation to risk
179
+ of nonfatal myocardial infarction and fatal CHD in the Nurses'' Health Study.
180
+ They assessed dietary information from the study''s 1990, 1994, and 1998 food
181
+ frequency questionnaires and computed cumulative average intakes of flavonols
182
+ and flavones. Cox proportional hazards regression with time-varying variables
183
+ was used for analysis. During 12 years of follow-up (1990-2002), the authors documented
184
+ 938 nonfatal myocardial infarctions and 324 CHD deaths among 66,360 women. They
185
+ observed no association between flavonol or flavone intake and risk of nonfatal
186
+ myocardial infarction or fatal CHD. However, a weak risk reduction for CHD death
187
+ was found among women with a higher intake of kaempferol, an individual flavonol
188
+ found primarily in broccoli and tea. Women in the highest quintile of kaempferol
189
+ intake relative to those in the lowest had a multivariate relative risk of 0.66
190
+ (95% confidence interval: 0.48, 0.93; p for trend = 0.04). The lower risk associated
191
+ with kaempferol intake was probably attributable to broccoli consumption. These
192
+ prospective data do not support an inverse association between flavonol or flavone
193
+ intake and CHD risk.'
194
+ - source_sentence: 'query: How Do Leafy Greens Protect Your Eyes?'
195
+ sentences:
196
+ - 'passsage: Increasing evidence suggests that acetaldehyde, the first and genotoxic
197
+ metabolite of ethanol, mediates the carcinogenicity of alcoholic beverages. Ethanol
198
+ is also contained in a number of ready-to-use mouthwashes typically between 5
199
+ and 27% vol. An increased risk of oral cancer has been discussed for users of
200
+ such mouthwashes; however, epidemiological evidence had remained inconclusive.
201
+ This study is the first to investigate acetaldehyde levels in saliva after use
202
+ of alcohol-containing mouthwashes. Ready-to-use mouthwashes and mouthrinses (n
203
+ = 13) were rinsed in the mouth by healthy, nonsmoking volunteers (n = 4) as intended
204
+ by the manufacturers (20 ml for 30 sec). Saliva was collected at 0.5, 2, 5 and
205
+ 10 min after mouthwash use and analyzed using headspace gas chromatography. The
206
+ acetaldehyde content in the saliva was 41 +/- 15 microM, range 9-85 microM (0.5
207
+ min), 52 +/- 14 microM, range 11-105 microM (2 min), 32 +/- 7 microM, range 9-67
208
+ microM (5 min) and 15 +/- 7 microM, range 0-37 microM (10 min). The contents were
209
+ significantly above endogenous levels and corresponding to concentrations normally
210
+ found after alcoholic beverage consumption. A twice-daily use of alcohol-containing
211
+ mouthwashes leads to a systemic acetaldehyde exposure of 0.26 microg/kg bodyweight/day
212
+ on average, which corresponds to a lifetime cancer risk of 3E-6. The margin of
213
+ exposure was calculated to be 217,604, which would be seen as a low public health
214
+ concern. However, the local acetaldehyde contents in the saliva are reaching concentrations
215
+ associated with DNA adduct formation and sister chromatid exchange in vitro, so
216
+ that concerns for local carcinogenic effects in the oral cavity remain.'
217
+ - 'passsage: Excitement about neurogenetics in the last two decades has diverted
218
+ attention from environmental causes of sporadic ALS. Fifty years ago endemic foci
219
+ of ALS with a frequency one hundred times that in the rest of the world attracted
220
+ attention since they offered the possibility of finding the cause for non-endemic
221
+ ALS throughout the world. Research on Guam suggested that ALS, Parkinson''s disease
222
+ and dementia (the ALS/PDC complex) was due to a neurotoxic non-protein amino acid,
223
+ beta-methylamino-L-alanine (BMAA), in the seeds of the cycad Cycas micronesica.
224
+ Recent discoveries that found that BMAA is produced by symbiotic cyanobacteria
225
+ within specialized roots of the cycads; that the concentration of protein-bound
226
+ BMAA is up to a hundred-fold greater than free BMAA in the seeds and flour; that
227
+ various animals forage on the seeds (flying foxes, pigs, deer), leading to biomagnification
228
+ up the food chain in Guam; and that protein-bound BMAA occurs in the brains of
229
+ Guamanians dying of ALS/PDC (average concentration 627 microg/g, 5 mM) but not
230
+ in control brains have rekindled interest in BMAA as a possible trigger for Guamanian
231
+ ALS/PDC. Perhaps most intriguing is the finding that BMAA is present in brain
232
+ tissues of North American patients who had died of Alzheimer''s disease (average
233
+ concentration 95 microg/g, 0.8mM); this suggests a possible etiological role for
234
+ BMAA in non-Guamanian neurodegenerative diseases. Cyanobacteria are ubiquitous
235
+ throughout the world, so it is possible that all humans are exposed to low amounts
236
+ of cyanobacterial BMAA, that protein-bound BMAA in human brains is a reservoir
237
+ for chronic neurotoxicity, and that cyanobacterial BMAA is a major cause of progressive
238
+ neurodegenerative diseases including ALS worldwide. Though Montine et al., using
239
+ different HPLC method and assay techniques from those used by Cox and colleagues,
240
+ were unable to reproduce the findings of Murch et al., Mash and colleagues using
241
+ the original techniques of Murch et al. have recently confirmed the presence of
242
+ protein-bound BMAA in the brains of North American patients dying with ALS and
243
+ Alzheimer''s disease (concentrations >100 microg/g) but not in the brains of non-neurological
244
+ controls or Huntington''s disease. We hypothesize that individuals who develop
245
+ neurodegenerations may have a genetic susceptibility because of inability to prevent
246
+ BMAA accumulation in brain proteins and that the particular pattern of neurodegeneration
247
+ that develops depends on the polygenic background of the individual.'
248
+ - 'passsage: PURPOSE: To explore the association between the consumption of fruits
249
+ and vegetables and the presence of glaucoma. DESIGN: Cross-sectional cohort study.
250
+ METHODS: In a sample of 1,155 women located in multiple centers in the United
251
+ States, glaucoma specialists diagnosed glaucoma in at least one eye by assessing
252
+ optic nerve head photographs and 76-point suprathreshold screening visual fields.
253
+ Consumption of fruits and vegetables was assessed using the Block Food Frequency
254
+ Questionnaire. The relationship between selected fruit and vegetable consumption
255
+ and glaucoma was investigated using adjusted logistic regression models. RESULTS:
256
+ Among 1,155 women, 95 (8.2%) were diagnosed with glaucoma. In adjusted analysis,
257
+ the odds of glaucoma risk were decreased by 69% (odds ratio [OR], 0.31; 95% confidence
258
+ interval [CI], 0.11 to 0.91) in women who consumed at least one serving per month
259
+ of green collards and kale compared with those who consumed fewer than one serving
260
+ per month, by 64% (OR, 0.36; 95% CI, 0.17 to 0.77) in women who consumed more
261
+ than two servings per week of carrots compared with those who consumed fewer than
262
+ one serving per week, and by 47% (OR, 0.53; 95% CI, 0.29 to 0.97) in women who
263
+ consumed at least one serving per week of canned or dried peaches compared with
264
+ those who consumed fewer than one serving per month. CONCLUSIONS: A higher intake
265
+ of certain fruits and vegetables may be associated with a decreased risk of glaucoma.
266
+ More studies are needed to investigate this relationship.'
267
+ - source_sentence: 'query: How Much Isoflavones Do I Need to Reduce Ovarian Cancer
268
+ Risk?'
269
+ sentences:
270
+ - 'passsage: It is likely that plant food consumption throughout much of human evolution
271
+ shaped the dietary requirements of contemporary humans. Diets would have been
272
+ high in dietary fiber, vegetable protein, plant sterols and associated phytochemicals,
273
+ and low in saturated and trans-fatty acids and other substrates for cholesterol
274
+ biosynthesis. To meet the body''s needs for cholesterol, we believe genetic differences
275
+ and polymorphisms were conserved by evolution, which tended to raise serum cholesterol
276
+ levels. As a result modern man, with a radically different diet and lifestyle,
277
+ especially in middle age, is now recommended to take medications to lower cholesterol
278
+ and reduce the risk of cardiovascular disease. Experimental introduction of high
279
+ intakes of viscous fibers, vegetable proteins and plant sterols in the form of
280
+ a possible Myocene diet of leafy vegetables, fruit and nuts, lowered serum LDL-cholesterol
281
+ in healthy volunteers by over 30%, equivalent to first generation statins, the
282
+ standard cholesterol-lowering medications. Furthermore, supplementation of a modern
283
+ therapeutic diet in hyperlipidemic subjects with the same components taken as
284
+ oat, barley and psyllium for viscous fibers, soy and almonds for vegetable proteins
285
+ and plant sterol-enriched margarine produced similar reductions in LDL-cholesterol
286
+ as the Myocene-like diet and reduced the majority of subjects'' blood lipids concentrations
287
+ into the normal range. We conclude that reintroduction of plant food components,
288
+ which would have been present in large quantities in the plant based diets eaten
289
+ throughout most of human evolution into modern diets can correct the lipid abnormalities
290
+ associated with contemporary eating patterns and reduce the need for pharmacological
291
+ interventions.'
292
+ - 'passsage: Liquid dietary supplements represent a fast growing market segment,
293
+ including botanically-based beverages containing mangosteen, acai, and noni. These
294
+ products often resemble fruit juice in packaging and appearance, but may contain
295
+ pharmacologically active ingredients. While little is known about the human health
296
+ effects or safety of consuming such products, manufacturers make extensive use
297
+ of low-quality published research to promote their products. This report analyzes
298
+ the science-based marketing claims of two of the most widely consumed mangosteen
299
+ liquid dietary supplements, and compares them to the findings of the research
300
+ being cited. The reviewer found that analyzed marketing claims overstate the significance
301
+ of findings, and fail to disclose severe methodological weaknesses of the research
302
+ they cite. If this trend extends to other related products that are similarly
303
+ widely consumed, it may pose a public health threat by misleading consumers into
304
+ assuming that product safety and effectiveness are backed by rigorous scientific
305
+ data.'
306
+ - 'passsage: Dietary phytochemical compounds, including isoflavones and isothiocyanates,
307
+ may inhibit cancer development but have not yet been examined in prospective epidemiologic
308
+ studies of ovarian cancer. The authors have investigated the association between
309
+ consumption of these and other nutrients and ovarian cancer risk in a prospective
310
+ cohort study. Among 97,275 eligible women in the California Teachers Study cohort
311
+ who completed the baseline dietary assessment in 1995–1996, 280 women developed
312
+ invasive or borderline ovarian cancer by December 31, 2003. Multivariable Cox
313
+ proportional hazards regression, with age as the timescale, was used to estimate
314
+ relative risks and 95% confidence intervals; all statistical tests were two sided.
315
+ Intake of isoflavones was associated with lower risk of ovarian cancer. Compared
316
+ with the risk for women who consumed less than 1 mg of total isoflavones per day,
317
+ the relative risk of ovarian cancer associated with consumption of more than 3
318
+ mg/day was 0.56 (95% confidence interval: 0.33, 0.96). Intake of isothiocyanates
319
+ or foods high in isothiocyanates was not associated with ovarian cancer risk,
320
+ nor was intake of macronutrients, antioxidant vitamins, or other micronutrients.
321
+ Although dietary consumption of isoflavones may be associated with decreased ovarian
322
+ cancer risk, most dietary factors are unlikely to play a major role in ovarian
323
+ cancer development.'
324
+ pipeline_tag: sentence-similarity
325
+ library_name: sentence-transformers
326
+ ---
327
+
328
+ # SentenceTransformer based on intfloat/e5-small-v2
329
+
330
+ This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [intfloat/e5-small-v2](https://huggingface.co/intfloat/e5-small-v2). It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
331
+
332
+ ## Model Details
333
+
334
+ ### Model Description
335
+ - **Model Type:** Sentence Transformer
336
+ - **Base model:** [intfloat/e5-small-v2](https://huggingface.co/intfloat/e5-small-v2) <!-- at revision ffb93f3bd4047442299a41ebb6fa998a38507c52 -->
337
+ - **Maximum Sequence Length:** 512 tokens
338
+ - **Output Dimensionality:** 384 dimensions
339
+ - **Similarity Function:** Cosine Similarity
340
+ <!-- - **Training Dataset:** Unknown -->
341
+ <!-- - **Language:** Unknown -->
342
+ <!-- - **License:** Unknown -->
343
+
344
+ ### Model Sources
345
+
346
+ - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
347
+ - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)
348
+ - **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers)
349
+
350
+ ### Full Model Architecture
351
+
352
+ ```
353
+ SentenceTransformer(
354
+ (0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'BertModel'})
355
+ (1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
356
+ )
357
+ ```
358
+
359
+ ## Usage
360
+
361
+ ### Direct Usage (Sentence Transformers)
362
+
363
+ First install the Sentence Transformers library:
364
+
365
+ ```bash
366
+ pip install -U sentence-transformers
367
+ ```
368
+
369
+ Then you can load this model and run inference.
370
+ ```python
371
+ from sentence_transformers import SentenceTransformer
372
+
373
+ # Download from the 🤗 Hub
374
+ model = SentenceTransformer("sentence_transformers_model_id")
375
+ # Run inference
376
+ sentences = [
377
+ 'query: How Much Isoflavones Do I Need to Reduce Ovarian Cancer Risk?',
378
+ 'passsage: Dietary phytochemical compounds, including isoflavones and isothiocyanates, may inhibit cancer development but have not yet been examined in prospective epidemiologic studies of ovarian cancer. The authors have investigated the association between consumption of these and other nutrients and ovarian cancer risk in a prospective cohort study. Among 97,275 eligible women in the California Teachers Study cohort who completed the baseline dietary assessment in 1995–1996, 280 women developed invasive or borderline ovarian cancer by December 31, 2003. Multivariable Cox proportional hazards regression, with age as the timescale, was used to estimate relative risks and 95% confidence intervals; all statistical tests were two sided. Intake of isoflavones was associated with lower risk of ovarian cancer. Compared with the risk for women who consumed less than 1 mg of total isoflavones per day, the relative risk of ovarian cancer associated with consumption of more than 3 mg/day was 0.56 (95% confidence interval: 0.33, 0.96). Intake of isothiocyanates or foods high in isothiocyanates was not associated with ovarian cancer risk, nor was intake of macronutrients, antioxidant vitamins, or other micronutrients. Although dietary consumption of isoflavones may be associated with decreased ovarian cancer risk, most dietary factors are unlikely to play a major role in ovarian cancer development.',
379
+ "passsage: It is likely that plant food consumption throughout much of human evolution shaped the dietary requirements of contemporary humans. Diets would have been high in dietary fiber, vegetable protein, plant sterols and associated phytochemicals, and low in saturated and trans-fatty acids and other substrates for cholesterol biosynthesis. To meet the body's needs for cholesterol, we believe genetic differences and polymorphisms were conserved by evolution, which tended to raise serum cholesterol levels. As a result modern man, with a radically different diet and lifestyle, especially in middle age, is now recommended to take medications to lower cholesterol and reduce the risk of cardiovascular disease. Experimental introduction of high intakes of viscous fibers, vegetable proteins and plant sterols in the form of a possible Myocene diet of leafy vegetables, fruit and nuts, lowered serum LDL-cholesterol in healthy volunteers by over 30%, equivalent to first generation statins, the standard cholesterol-lowering medications. Furthermore, supplementation of a modern therapeutic diet in hyperlipidemic subjects with the same components taken as oat, barley and psyllium for viscous fibers, soy and almonds for vegetable proteins and plant sterol-enriched margarine produced similar reductions in LDL-cholesterol as the Myocene-like diet and reduced the majority of subjects' blood lipids concentrations into the normal range. We conclude that reintroduction of plant food components, which would have been present in large quantities in the plant based diets eaten throughout most of human evolution into modern diets can correct the lipid abnormalities associated with contemporary eating patterns and reduce the need for pharmacological interventions.",
380
+ ]
381
+ embeddings = model.encode(sentences)
382
+ print(embeddings.shape)
383
+ # [3, 384]
384
+
385
+ # Get the similarity scores for the embeddings
386
+ similarities = model.similarity(embeddings, embeddings)
387
+ print(similarities)
388
+ # tensor([[ 1.0000, 0.8492, 0.0442],
389
+ # [ 0.8492, 1.0000, -0.0126],
390
+ # [ 0.0442, -0.0126, 1.0000]])
391
+ ```
392
+
393
+ <!--
394
+ ### Direct Usage (Transformers)
395
+
396
+ <details><summary>Click to see the direct usage in Transformers</summary>
397
+
398
+ </details>
399
+ -->
400
+
401
+ <!--
402
+ ### Downstream Usage (Sentence Transformers)
403
+
404
+ You can finetune this model on your own dataset.
405
+
406
+ <details><summary>Click to expand</summary>
407
+
408
+ </details>
409
+ -->
410
+
411
+ <!--
412
+ ### Out-of-Scope Use
413
+
414
+ *List how the model may foreseeably be misused and address what users ought not to do with the model.*
415
+ -->
416
+
417
+ <!--
418
+ ## Bias, Risks and Limitations
419
+
420
+ *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
421
+ -->
422
+
423
+ <!--
424
+ ### Recommendations
425
+
426
+ *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
427
+ -->
428
+
429
+ ## Training Details
430
+
431
+ ### Training Dataset
432
+
433
+ #### Unnamed Dataset
434
+
435
+ * Size: 26,136 training samples
436
+ * Columns: <code>anchor</code> and <code>positive</code>
437
+ * Approximate statistics based on the first 1000 samples:
438
+ | | anchor | positive |
439
+ |:--------|:---------------------------------------------------------------------------------|:------------------------------------------------------------------------------------|
440
+ | type | string | string |
441
+ | details | <ul><li>min: 7 tokens</li><li>mean: 12.9 tokens</li><li>max: 27 tokens</li></ul> | <ul><li>min: 25 tokens</li><li>mean: 331.4 tokens</li><li>max: 512 tokens</li></ul> |
442
+ * Samples:
443
+ | anchor | positive |
444
+ |:------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
445
+ | <code>query: Does Turmeric Help With Ulcers?</code> | <code>passsage: The purpose of this review is to summarize the pertinent literature published in the present era regarding the antiulcerogenic property of curcumin against the pathological changes in response to ulcer effectors (Helicobacter pylori infection, chronic ingestion of non-steroidal anti-inflammatory drugs, and exogenous substances). The gastrointestinal problems caused by different etiologies was observed to be associated with the alterations of various physiologic parameters such as reactive oxygen species, nitric oxide synthase, lipid peroxidation, and secretion of excessive gastric acid. Gastrointestinal ulcer results probably due to imbalance between the aggressive and the defensive factors. In 80% of the cases, gastric ulcer is caused primarily due to the use of non-steroidal anti-inflammatory category of drug, 10% by H. pylori, and about 8-10% by the intake of very spicy and fast food. Although a number of antiulcer drugs and cytoprotectants are available, all these drugs h...</code> |
446
+ | <code>query: What Are the Dangers of Environmental Toxins for Men's Reproductive Health?</code> | <code>passsage: Male reproductive disorders that are of interest from an environmental point of view include sexual dysfunction, infertility, cryptorchidism, hypospadias and testicular cancer. Several reports suggest declining sperm counts and increase of these reproductive disorders in some areas during some time periods past 50 years. Except for testicular cancer this evidence is circumstantial and needs cautious interpretation. However, the male germ line is one of the most sensitive tissues to the damaging effects of ionizing radiation, radiant heat and a number of known toxicants. So far occupational hazards are the best documented risk factors for impaired male reproductive function and include physical exposures (radiant heat, ionizing radiation, high frequency electromagnetic radiation), chemical exposures (some solvents as carbon disulfide and ethylene glycol ethers, some pesticides as dibromochloropropane, ethylendibromide and DDT/DDE, some heavy metals as inorganic lead and mercur...</code> |
447
+ | <code>query: Does Flaxseed Help Lower Blood Pressure?</code> | <code>passsage: Flaxseed contains ω-3 fatty acids, lignans, and fiber that together may provide benefits to patients with cardiovascular disease. Animal work identified that patients with peripheral artery disease may particularly benefit from dietary supplementation with flaxseed. Hypertension is commonly associated with peripheral artery disease. The purpose of the study was to examine the effects of daily ingestion of flaxseed on systolic (SBP) and diastolic blood pressure (DBP) in peripheral artery disease patients. In this prospective, double-blinded, placebo-controlled, randomized trial, patients (110 in total) ingested a variety of foods that contained 30 g of milled flaxseed or placebo each day over 6 months. Plasma levels of the ω-3 fatty acid α-linolenic acid and enterolignans increased 2- to 50-fold in the flaxseed-fed group but did not increase significantly in the placebo group. Patient body weights were not significantly different between the 2 groups at any time. SBP was ≈ 10 ...</code> |
448
+ * Loss: [<code>CachedMultipleNegativesRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with these parameters:
449
+ ```json
450
+ {
451
+ "scale": 20.0,
452
+ "similarity_fct": "cos_sim",
453
+ "mini_batch_size": 16,
454
+ "gather_across_devices": false
455
+ }
456
+ ```
457
+
458
+ ### Evaluation Dataset
459
+
460
+ #### Unnamed Dataset
461
+
462
+ * Size: 2,904 evaluation samples
463
+ * Columns: <code>anchor</code> and <code>positive</code>
464
+ * Approximate statistics based on the first 1000 samples:
465
+ | | anchor | positive |
466
+ |:--------|:----------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|
467
+ | type | string | string |
468
+ | details | <ul><li>min: 7 tokens</li><li>mean: 13.08 tokens</li><li>max: 27 tokens</li></ul> | <ul><li>min: 23 tokens</li><li>mean: 335.36 tokens</li><li>max: 512 tokens</li></ul> |
469
+ * Samples:
470
+ | anchor | positive |
471
+ |:-------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
472
+ | <code>query: Does Diet Reduce Alzheimer's Risk?</code> | <code>passsage: BACKGROUND: Numerous studies have investigated risk factors for Alzheimer disease (AD). However, at a recent National Institutes of Health State-of-the-Science Conference, an independent panel found insufficient evidence to support the association of any modifiable factor with risk of cognitive decline or AD. OBJECTIVE: To present key findings for selected factors and AD risk that led the panel to their conclusion. DATA SOURCES: An evidence report was commissioned by the Agency for Healthcare Research and Quality. It included English-language publications in MEDLINE and the Cochrane Database of Systematic Reviews from 1984 through October 27, 2009. Expert presentations and public discussions were considered. STUDY SELECTION: Study inclusion criteria for the evidence report were participants aged 50 years and older from general populations in developed countries; minimum sample sizes of 300 for cohort studies and 50 for randomized controlled trials; at least 2 years between ex...</code> |
473
+ | <code>query: Is Spam Bad For Your Diabetes?</code> | <code>passsage: Background: Fifty percent of American Indians (AIs) develop diabetes by age 55 y. Whether processed meat is associated with the risk of diabetes in AIs, a rural population with a high intake of processed meat (eg, canned meats in general, referred to as “spam”) and a high rate of diabetes, is unknown. Objective: We examined the associations of usual intake of processed meat with incident diabetes in AIs. Design: This prospective cohort study included AI participants from the Strong Heart Family Study who were free of diabetes and cardiovascular disease at baseline and who participated in a 5-y follow-up examination (n = 2001). Dietary intake was ascertained by using a Block food-frequency questionnaire at baseline. Incident diabetes was defined on the basis of 2003 American Diabetes Association criteria. Generalized estimating equations were used to examine the associations of dietary intake with incident diabetes. Results: We identified 243 incident cases of diabetes. In a c...</code> |
474
+ | <code>query: Is Vitamin D Good For Cancer Prevention?</code> | <code>passsage: Observational and ecological studies are generally used to determine the presence of effect of cancer risk-modifying factors. Researchers generally agree that environmental factors such as smoking, alcohol consumption, poor diet, lack of physical activity, and low serum 25-hdyroxyvitamin D levels are important cancer risk factors. This ecological study used age-adjusted incidence rates for 21 cancers for 157 countries (87 with high-quality data) in 2008 with respect to dietary supply and other factors, including per capita gross domestic product, life expectancy, lung cancer incidence rate (an index for smoking), and latitude (an index for solar ultraviolet-B doses). The factors found to correlate strongly with multiple types of cancer were lung cancer (direct correlation with 12 types of cancer), energy derived from animal products (direct correlation with 12 types of cancer, inverse with two), latitude (direct correlation with six types, inverse correlation with three), and...</code> |
475
+ * Loss: [<code>CachedMultipleNegativesRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with these parameters:
476
+ ```json
477
+ {
478
+ "scale": 20.0,
479
+ "similarity_fct": "cos_sim",
480
+ "mini_batch_size": 16,
481
+ "gather_across_devices": false
482
+ }
483
+ ```
484
+
485
+ ### Training Hyperparameters
486
+ #### Non-Default Hyperparameters
487
+
488
+ - `eval_strategy`: epoch
489
+ - `per_device_train_batch_size`: 128
490
+ - `learning_rate`: 2e-05
491
+ - `num_train_epochs`: 30
492
+ - `warmup_ratio`: 0.1
493
+ - `fp16`: True
494
+ - `load_best_model_at_end`: True
495
+ - `batch_sampler`: no_duplicates
496
+
497
+ #### All Hyperparameters
498
+ <details><summary>Click to expand</summary>
499
+
500
+ - `overwrite_output_dir`: False
501
+ - `do_predict`: False
502
+ - `eval_strategy`: epoch
503
+ - `prediction_loss_only`: True
504
+ - `per_device_train_batch_size`: 128
505
+ - `per_device_eval_batch_size`: 8
506
+ - `per_gpu_train_batch_size`: None
507
+ - `per_gpu_eval_batch_size`: None
508
+ - `gradient_accumulation_steps`: 1
509
+ - `eval_accumulation_steps`: None
510
+ - `torch_empty_cache_steps`: None
511
+ - `learning_rate`: 2e-05
512
+ - `weight_decay`: 0.0
513
+ - `adam_beta1`: 0.9
514
+ - `adam_beta2`: 0.999
515
+ - `adam_epsilon`: 1e-08
516
+ - `max_grad_norm`: 1.0
517
+ - `num_train_epochs`: 30
518
+ - `max_steps`: -1
519
+ - `lr_scheduler_type`: linear
520
+ - `lr_scheduler_kwargs`: {}
521
+ - `warmup_ratio`: 0.1
522
+ - `warmup_steps`: 0
523
+ - `log_level`: passive
524
+ - `log_level_replica`: warning
525
+ - `log_on_each_node`: True
526
+ - `logging_nan_inf_filter`: True
527
+ - `save_safetensors`: True
528
+ - `save_on_each_node`: False
529
+ - `save_only_model`: False
530
+ - `restore_callback_states_from_checkpoint`: False
531
+ - `no_cuda`: False
532
+ - `use_cpu`: False
533
+ - `use_mps_device`: False
534
+ - `seed`: 42
535
+ - `data_seed`: None
536
+ - `jit_mode_eval`: False
537
+ - `bf16`: False
538
+ - `fp16`: True
539
+ - `fp16_opt_level`: O1
540
+ - `half_precision_backend`: auto
541
+ - `bf16_full_eval`: False
542
+ - `fp16_full_eval`: False
543
+ - `tf32`: None
544
+ - `local_rank`: 0
545
+ - `ddp_backend`: None
546
+ - `tpu_num_cores`: None
547
+ - `tpu_metrics_debug`: False
548
+ - `debug`: []
549
+ - `dataloader_drop_last`: False
550
+ - `dataloader_num_workers`: 0
551
+ - `dataloader_prefetch_factor`: None
552
+ - `past_index`: -1
553
+ - `disable_tqdm`: False
554
+ - `remove_unused_columns`: True
555
+ - `label_names`: None
556
+ - `load_best_model_at_end`: True
557
+ - `ignore_data_skip`: False
558
+ - `fsdp`: []
559
+ - `fsdp_min_num_params`: 0
560
+ - `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
561
+ - `fsdp_transformer_layer_cls_to_wrap`: None
562
+ - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
563
+ - `parallelism_config`: None
564
+ - `deepspeed`: None
565
+ - `label_smoothing_factor`: 0.0
566
+ - `optim`: adamw_torch_fused
567
+ - `optim_args`: None
568
+ - `adafactor`: False
569
+ - `group_by_length`: False
570
+ - `length_column_name`: length
571
+ - `project`: huggingface
572
+ - `trackio_space_id`: trackio
573
+ - `ddp_find_unused_parameters`: None
574
+ - `ddp_bucket_cap_mb`: None
575
+ - `ddp_broadcast_buffers`: False
576
+ - `dataloader_pin_memory`: True
577
+ - `dataloader_persistent_workers`: False
578
+ - `skip_memory_metrics`: True
579
+ - `use_legacy_prediction_loop`: False
580
+ - `push_to_hub`: False
581
+ - `resume_from_checkpoint`: None
582
+ - `hub_model_id`: None
583
+ - `hub_strategy`: every_save
584
+ - `hub_private_repo`: None
585
+ - `hub_always_push`: False
586
+ - `hub_revision`: None
587
+ - `gradient_checkpointing`: False
588
+ - `gradient_checkpointing_kwargs`: None
589
+ - `include_inputs_for_metrics`: False
590
+ - `include_for_metrics`: []
591
+ - `eval_do_concat_batches`: True
592
+ - `fp16_backend`: auto
593
+ - `push_to_hub_model_id`: None
594
+ - `push_to_hub_organization`: None
595
+ - `mp_parameters`:
596
+ - `auto_find_batch_size`: False
597
+ - `full_determinism`: False
598
+ - `torchdynamo`: None
599
+ - `ray_scope`: last
600
+ - `ddp_timeout`: 1800
601
+ - `torch_compile`: False
602
+ - `torch_compile_backend`: None
603
+ - `torch_compile_mode`: None
604
+ - `include_tokens_per_second`: False
605
+ - `include_num_input_tokens_seen`: no
606
+ - `neftune_noise_alpha`: None
607
+ - `optim_target_modules`: None
608
+ - `batch_eval_metrics`: False
609
+ - `eval_on_start`: False
610
+ - `use_liger_kernel`: False
611
+ - `liger_kernel_config`: None
612
+ - `eval_use_gather_object`: False
613
+ - `average_tokens_across_devices`: True
614
+ - `prompts`: None
615
+ - `batch_sampler`: no_duplicates
616
+ - `multi_dataset_batch_sampler`: proportional
617
+ - `router_mapping`: {}
618
+ - `learning_rate_mapping`: {}
619
+
620
+ </details>
621
+
622
+ ### Training Logs
623
+ | Epoch | Step | Training Loss | Validation Loss |
624
+ |:--------:|:--------:|:-------------:|:---------------:|
625
+ | 1.0 | 205 | 1.4561 | 0.0201 |
626
+ | 2.0 | 410 | 0.1947 | 0.0127 |
627
+ | 3.0 | 615 | 0.1403 | 0.0114 |
628
+ | 4.0 | 820 | 0.1137 | 0.0088 |
629
+ | 5.0 | 1025 | 0.0911 | 0.0080 |
630
+ | 6.0 | 1230 | 0.0818 | 0.0084 |
631
+ | 7.0 | 1435 | 0.0734 | 0.0078 |
632
+ | 8.0 | 1640 | 0.067 | 0.0076 |
633
+ | 9.0 | 1845 | 0.0607 | 0.0082 |
634
+ | 10.0 | 2050 | 0.0553 | 0.0073 |
635
+ | 11.0 | 2255 | 0.0526 | 0.0065 |
636
+ | 12.0 | 2460 | 0.0505 | 0.0066 |
637
+ | 13.0 | 2665 | 0.0505 | 0.0063 |
638
+ | 14.0 | 2870 | 0.0466 | 0.0060 |
639
+ | 15.0 | 3075 | 0.0448 | 0.0063 |
640
+ | 16.0 | 3280 | 0.045 | 0.0060 |
641
+ | 17.0 | 3485 | 0.0417 | 0.0057 |
642
+ | 18.0 | 3690 | 0.0395 | 0.0057 |
643
+ | 19.0 | 3895 | 0.0425 | 0.0065 |
644
+ | 20.0 | 4100 | 0.0379 | 0.0058 |
645
+ | 21.0 | 4305 | 0.0389 | 0.0055 |
646
+ | 22.0 | 4510 | 0.0394 | 0.0055 |
647
+ | 23.0 | 4715 | 0.0351 | 0.0057 |
648
+ | 24.0 | 4920 | 0.0359 | 0.0054 |
649
+ | **25.0** | **5125** | **0.0366** | **0.0054** |
650
+ | 26.0 | 5330 | 0.0356 | 0.0055 |
651
+ | 27.0 | 5535 | 0.0347 | 0.0055 |
652
+ | 28.0 | 5740 | 0.0346 | 0.0054 |
653
+
654
+ * The bold row denotes the saved checkpoint.
655
+
656
+ ### Framework Versions
657
+ - Python: 3.12.12
658
+ - Sentence Transformers: 5.1.2
659
+ - Transformers: 4.57.1
660
+ - PyTorch: 2.8.0+cu126
661
+ - Accelerate: 1.11.0
662
+ - Datasets: 4.0.0
663
+ - Tokenizers: 0.22.1
664
+
665
+ ## Citation
666
+
667
+ ### BibTeX
668
+
669
+ #### Sentence Transformers
670
+ ```bibtex
671
+ @inproceedings{reimers-2019-sentence-bert,
672
+ title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
673
+ author = "Reimers, Nils and Gurevych, Iryna",
674
+ booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
675
+ month = "11",
676
+ year = "2019",
677
+ publisher = "Association for Computational Linguistics",
678
+ url = "https://arxiv.org/abs/1908.10084",
679
+ }
680
+ ```
681
+
682
+ #### CachedMultipleNegativesRankingLoss
683
+ ```bibtex
684
+ @misc{gao2021scaling,
685
+ title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
686
+ author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
687
+ year={2021},
688
+ eprint={2101.06983},
689
+ archivePrefix={arXiv},
690
+ primaryClass={cs.LG}
691
+ }
692
+ ```
693
+
694
+ <!--
695
+ ## Glossary
696
+
697
+ *Clearly define terms in order to be accessible across audiences.*
698
+ -->
699
+
700
+ <!--
701
+ ## Model Card Authors
702
+
703
+ *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
704
+ -->
705
+
706
+ <!--
707
+ ## Model Card Contact
708
+
709
+ *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
710
+ -->
config.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "BertModel"
4
+ ],
5
+ "attention_probs_dropout_prob": 0.1,
6
+ "classifier_dropout": null,
7
+ "dtype": "float32",
8
+ "hidden_act": "gelu",
9
+ "hidden_dropout_prob": 0.1,
10
+ "hidden_size": 384,
11
+ "initializer_range": 0.02,
12
+ "intermediate_size": 1536,
13
+ "layer_norm_eps": 1e-12,
14
+ "max_position_embeddings": 512,
15
+ "model_type": "bert",
16
+ "num_attention_heads": 12,
17
+ "num_hidden_layers": 12,
18
+ "pad_token_id": 0,
19
+ "position_embedding_type": "absolute",
20
+ "transformers_version": "4.57.1",
21
+ "type_vocab_size": 2,
22
+ "use_cache": true,
23
+ "vocab_size": 30522
24
+ }
config_sentence_transformers.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "SentenceTransformer",
3
+ "__version__": {
4
+ "sentence_transformers": "5.1.2",
5
+ "transformers": "4.57.1",
6
+ "pytorch": "2.8.0+cu126"
7
+ },
8
+ "prompts": {
9
+ "query": "",
10
+ "document": ""
11
+ },
12
+ "default_prompt_name": null,
13
+ "similarity_fn_name": "cosine"
14
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b25b830ec9c929c1f7354cc2aa8f547ee8621f84dbfc535e3038b249ae4d6018
3
+ size 133462128
modules.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "idx": 0,
4
+ "name": "0",
5
+ "path": "",
6
+ "type": "sentence_transformers.models.Transformer"
7
+ },
8
+ {
9
+ "idx": 1,
10
+ "name": "1",
11
+ "path": "1_Pooling",
12
+ "type": "sentence_transformers.models.Pooling"
13
+ }
14
+ ]
runs/Nov02_13-44-08_8861e3a0a32b/events.out.tfevents.1762091052.8861e3a0a32b.3168.0 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:559455eb346030078a0f770ed404a96cbefc6a8ca60c57a0dbd8f315e220047d
3
+ size 18475
sentence_bert_config.json ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {
2
+ "max_seq_length": 512,
3
+ "do_lower_case": false
4
+ }
special_tokens_map.json ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "cls_token": {
3
+ "content": "[CLS]",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "mask_token": {
10
+ "content": "[MASK]",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "pad_token": {
17
+ "content": "[PAD]",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ },
23
+ "sep_token": {
24
+ "content": "[SEP]",
25
+ "lstrip": false,
26
+ "normalized": false,
27
+ "rstrip": false,
28
+ "single_word": false
29
+ },
30
+ "unk_token": {
31
+ "content": "[UNK]",
32
+ "lstrip": false,
33
+ "normalized": false,
34
+ "rstrip": false,
35
+ "single_word": false
36
+ }
37
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "added_tokens_decoder": {
3
+ "0": {
4
+ "content": "[PAD]",
5
+ "lstrip": false,
6
+ "normalized": false,
7
+ "rstrip": false,
8
+ "single_word": false,
9
+ "special": true
10
+ },
11
+ "100": {
12
+ "content": "[UNK]",
13
+ "lstrip": false,
14
+ "normalized": false,
15
+ "rstrip": false,
16
+ "single_word": false,
17
+ "special": true
18
+ },
19
+ "101": {
20
+ "content": "[CLS]",
21
+ "lstrip": false,
22
+ "normalized": false,
23
+ "rstrip": false,
24
+ "single_word": false,
25
+ "special": true
26
+ },
27
+ "102": {
28
+ "content": "[SEP]",
29
+ "lstrip": false,
30
+ "normalized": false,
31
+ "rstrip": false,
32
+ "single_word": false,
33
+ "special": true
34
+ },
35
+ "103": {
36
+ "content": "[MASK]",
37
+ "lstrip": false,
38
+ "normalized": false,
39
+ "rstrip": false,
40
+ "single_word": false,
41
+ "special": true
42
+ }
43
+ },
44
+ "clean_up_tokenization_spaces": true,
45
+ "cls_token": "[CLS]",
46
+ "do_lower_case": true,
47
+ "extra_special_tokens": {},
48
+ "mask_token": "[MASK]",
49
+ "model_max_length": 512,
50
+ "pad_token": "[PAD]",
51
+ "sep_token": "[SEP]",
52
+ "strip_accents": null,
53
+ "tokenize_chinese_chars": true,
54
+ "tokenizer_class": "BertTokenizer",
55
+ "unk_token": "[UNK]"
56
+ }
training_args.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f659079ee3e5f813861ac2e046557045bc0cc4e8371b4018b1d2e789122e5676
3
+ size 6289
vocab.txt ADDED
The diff for this file is too large to render. See raw diff