<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
		<id>https://www.scipedia.com/wd/index.php?action=history&amp;feed=atom&amp;title=Dugre_et_al_2019a</id>
		<title>Dugre et al 2019a - Revision history</title>
		<link rel="self" type="application/atom+xml" href="https://www.scipedia.com/wd/index.php?action=history&amp;feed=atom&amp;title=Dugre_et_al_2019a"/>
		<link rel="alternate" type="text/html" href="https://www.scipedia.com/wd/index.php?title=Dugre_et_al_2019a&amp;action=history"/>
		<updated>2026-04-17T06:03:36Z</updated>
		<subtitle>Revision history for this page on the wiki</subtitle>
		<generator>MediaWiki 1.27.0-wmf.10</generator>

	<entry>
		<id>https://www.scipedia.com/wd/index.php?title=Dugre_et_al_2019a&amp;diff=191621&amp;oldid=prev</id>
		<title>Scipediacontent: Scipediacontent moved page Draft Content 481826966 to Dugre et al 2019a</title>
		<link rel="alternate" type="text/html" href="https://www.scipedia.com/wd/index.php?title=Dugre_et_al_2019a&amp;diff=191621&amp;oldid=prev"/>
				<updated>2021-01-28T16:45:32Z</updated>
		
		<summary type="html">&lt;p&gt;Scipediacontent moved page &lt;a href=&quot;/public/Draft_Content_481826966&quot; class=&quot;mw-redirect&quot; title=&quot;Draft Content 481826966&quot;&gt;Draft Content 481826966&lt;/a&gt; to &lt;a href=&quot;/public/Dugre_et_al_2019a&quot; title=&quot;Dugre et al 2019a&quot;&gt;Dugre et al 2019a&lt;/a&gt;&lt;/p&gt;
&lt;table class=&quot;diff diff-contentalign-left&quot; data-mw=&quot;interface&quot;&gt;
				&lt;tr style='vertical-align: top;' lang='en'&gt;
				&lt;td colspan='1' style=&quot;background-color: white; color:black; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan='1' style=&quot;background-color: white; color:black; text-align: center;&quot;&gt;Revision as of 16:45, 28 January 2021&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan='2' style='text-align: center;' lang='en'&gt;&lt;div class=&quot;mw-diff-empty&quot;&gt;(No difference)&lt;/div&gt;
&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;</summary>
		<author><name>Scipediacontent</name></author>	</entry>

	<entry>
		<id>https://www.scipedia.com/wd/index.php?title=Dugre_et_al_2019a&amp;diff=191620&amp;oldid=prev</id>
		<title>Scipediacontent: Created page with &quot; == Abstract ==  In the past few years, neuroimaging has entered the Big Data era due to the joint increase in image resolution, data sharing, and study sizes. However, no par...&quot;</title>
		<link rel="alternate" type="text/html" href="https://www.scipedia.com/wd/index.php?title=Dugre_et_al_2019a&amp;diff=191620&amp;oldid=prev"/>
				<updated>2021-01-28T16:45:29Z</updated>
		
		<summary type="html">&lt;p&gt;Created page with &amp;quot; == Abstract ==  In the past few years, neuroimaging has entered the Big Data era due to the joint increase in image resolution, data sharing, and study sizes. However, no par...&amp;quot;&lt;/p&gt;
&lt;p&gt;&lt;b&gt;New page&lt;/b&gt;&lt;/p&gt;&lt;div&gt;&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
In the past few years, neuroimaging has entered the Big Data era due to the joint increase in image resolution, data sharing, and study sizes. However, no particular Big Data engines have emerged in this field, and several alternatives remain available. We compare two popular Big Data engines with Python APIs, Apache Spark and Dask, for their runtime performance in processing neuroimaging pipelines. Our evaluation uses two synthetic pipelines processing the 81GB BigBrain image, and a real pipeline processing anatomical data from more than 1,000 subjects. We benchmark these pipelines using various combinations of task durations, data sizes, and numbers of workers, deployed on an 8-node (8 cores ea.) compute cluster in Compute Canada's Arbutus cloud. We evaluate PySpark's RDD API against Dask's Bag, Delayed and Futures. Results show that despite slight differences between Spark and Dask, both engines perform comparably. However, Dask pipelines risk being limited by Python's GIL depending on task type and cluster configuration. In all cases, the major limiting factor was data transfer. While either engine is suitable for neuroimaging pipelines, more effort needs to be placed in reducing data transfer time.&lt;br /&gt;
&lt;br /&gt;
Comment: 10 pages, 15 figures, 1 tables. To appear in the proceeding of the 14th WORKS Workshop on Topics in Workflows in Support of Large-Scale Science, 17 November 2019, Denver, CO, USA&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Original document ==&lt;br /&gt;
&lt;br /&gt;
The different versions of the original document can be found in:&lt;br /&gt;
&lt;br /&gt;
* [http://arxiv.org/abs/1907.13030 http://arxiv.org/abs/1907.13030]&lt;br /&gt;
&lt;br /&gt;
* [http://arxiv.org/pdf/1907.13030 http://arxiv.org/pdf/1907.13030]&lt;br /&gt;
&lt;br /&gt;
* [http://xplorestaging.ieee.org/ielx7/8937308/8943488/08943502.pdf?arnumber=8943502 http://xplorestaging.ieee.org/ielx7/8937308/8943488/08943502.pdf?arnumber=8943502],&lt;br /&gt;
: [http://dx.doi.org/10.1109/works49585.2019.00010 http://dx.doi.org/10.1109/works49585.2019.00010]&lt;br /&gt;
&lt;br /&gt;
* [https://dblp.uni-trier.de/db/journals/corr/corr1907.html#abs-1907-13030 https://dblp.uni-trier.de/db/journals/corr/corr1907.html#abs-1907-13030],&lt;br /&gt;
: [https://arxiv.org/pdf/1907.13030.pdf https://arxiv.org/pdf/1907.13030.pdf],&lt;br /&gt;
: [https://arxiv.org/abs/1907.13030 https://arxiv.org/abs/1907.13030],&lt;br /&gt;
: [https://ui.adsabs.harvard.edu/abs/2019arXiv190713030D/abstract https://ui.adsabs.harvard.edu/abs/2019arXiv190713030D/abstract],&lt;br /&gt;
: [https://academic.microsoft.com/#/detail/3000137827 https://academic.microsoft.com/#/detail/3000137827]&lt;/div&gt;</summary>
		<author><name>Scipediacontent</name></author>	</entry>

	</feed>