<?xml version="1.0" encoding="utf-8"?>
<?xml-stylesheet type="text/xsl" href="../assets/xml/rss.xsl" media="all"?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>I Love Symposia! (Posts about science)</title><link>https://ilovesymposia.com/</link><description></description><atom:link href="https://ilovesymposia.com/categories/science.xml" rel="self" type="application/rss+xml"></atom:link><language>en</language><copyright>Contents © 2019 &lt;a href="mailto:jni.soma@fastmail.com"&gt;Juan Nunez-Iglesias&lt;/a&gt; 
&lt;a rel="license" href="https://creativecommons.org/licenses/by/4.0/"&gt;
&lt;img alt="Creative Commons License BY"
style="border-width:0; margin-bottom:12px;"
src="https://i.creativecommons.org/l/by/4.0/88x31.png"&gt;&lt;/a&gt;</copyright><lastBuildDate>Thu, 24 Oct 2019 00:09:12 GMT</lastBuildDate><generator>Nikola (getnikola.com)</generator><docs>http://blogs.law.harvard.edu/tech/rss</docs><item><title>Why citations are not enough for open source software</title><link>https://ilovesymposia.com/2019/05/28/why-citations-are-not-enough-for-open-source-software/</link><dc:creator>Juan Nunez-Iglesias</dc:creator><description>&lt;div&gt;&lt;p&gt;A few weeks ago I wrote about &lt;a href="https://ilovesymposia.com/2019/05/02/why-you-should-cite-open-source-tools/"&gt;why you should cite open source
tools&lt;/a&gt;.
Although I think citations important, though, there are major problems in
relying on them &lt;em&gt;alone&lt;/em&gt; to support open source work.&lt;/p&gt;
&lt;p&gt;The biggest problem is that papers describing a software library can only give
credit to the contributors at the time that the paper was written. The
preferred citation for the SciPy library is “Eric Jones, Travis Oliphant, Pearu
Peterson, &lt;em&gt;et al&lt;/em&gt;”, 2001. The “&lt;em&gt;et al&lt;/em&gt;” is not an abbreviation here, but a
fixed shorthand for all other contributors. Needless to say many, &lt;em&gt;many&lt;/em&gt; people
have contributed to the SciPy library since 2001 (GitHub counts 716
contributors as of this writing), and they are unable to get credit within the
academic system for those contributions. (As an aside, Google counts about
1,200 citations to SciPy, which is a breathtaking undercounting of its value
and influence, and reinforces my earlier point: &lt;strong&gt;cite open source software!
Definitely don't use this post as an excuse not to cite it!!!&lt;/strong&gt;)&lt;/p&gt;
&lt;p&gt;Not surprisingly, we have had massive contributions to scikit-image since our
2014 paper, and those contributors miss out on the citations to our paper.&lt;/p&gt;
&lt;!-- TEASER_END --&gt;

&lt;p&gt;From a maintainer point of view, one way to deal with this is to periodically
write “update” publications that become the preferred way to cite the
software. This is the approach taken by the
&lt;a href="https://cellprofiler.org/citations/"&gt;CellProfiler&lt;/a&gt; team, for example, and the
one we will take for scikit-image. This allows new contributors to benefit,
while also preventing infinite returns for the original authors, which is as
it should be.&lt;/p&gt;
&lt;p&gt;I think this is as good as project maintainers can do now, but leaves open the
questions of authorship, publication frequency, and the potential dilution of
the authors' message.&lt;/p&gt;
&lt;p&gt;A &lt;em&gt;proper&lt;/em&gt; fix requires structural change that directly recognises the value of
open source software to science. Software papers are in fact a &lt;em&gt;hack&lt;/em&gt; of the
academic system, a currency conversion, and as with any conversion, there are
losses and inefficiencies involved. Indeed, the paper provides less value than
good documentation for the software. (Last year, when someone tweeted about their
latest software paper, many replies asked for the GitHub link in the
abstract! People wanted to look at the software, not the paper.)&lt;/p&gt;
&lt;p&gt;So what does open source software &lt;em&gt;really&lt;/em&gt; need, if not citations? Of course,
it's money. Citations are useful because they can &lt;em&gt;potentially&lt;/em&gt; be translated
into grants, promotions, and jobs, but wouldn't it be great if we could cut out
the middleman and just have great open source translate &lt;em&gt;directly&lt;/em&gt; into
grants, promotions, and jobs?&lt;/p&gt;
&lt;p&gt;The insane thing is that granting bodies actually give enormous amounts of
money to &lt;em&gt;closed source&lt;/em&gt; software developers, since grant money is routinely
spent on software licenses. Instead of subsidising proprietary software, that
money could be used to fund open source developers, and improve software for
&lt;em&gt;everyone&lt;/em&gt;.&lt;/p&gt;
&lt;h3&gt;How to fund open source&lt;/h3&gt;
&lt;p&gt;I might have left the how for a later post, but while I was drafting this, CZI
&lt;a href="https://twitter.com/cziscience/status/1128693937130991623"&gt;announced&lt;/a&gt; direct
funding specifically for open source software. Needless to say I think this is
wonderful. I am funded by an earlier closed
&lt;a href="https://www.chanzuckerberg.com/newsroom/czi-announces-support-for-open-source-software-efforts-to-improve-biomedical-imaging"&gt;funding round&lt;/a&gt;
to develop scikit-image, and I thought a lot at the time about all the other
deserving, unfunded projects that I use. An open call is absolutely the right
thing to do, and supporting &lt;em&gt;maintenance&lt;/em&gt; rather than development of hot new
things, as this CZI call is doing, is also the right thing to do. (Others have
flagged that a single year of support, even if renewable, makes it difficult to
support careers, and I agree, but this is a new space, and it makes sense that
CZI would dip its toes before jmuping in.)&lt;/p&gt;
&lt;p&gt;The response to CZI's announcement has been absolute fire. There is enormous
pent-up need and
&lt;a href="https://twitter.com/amuellerml/status/1117455802598662144"&gt;frustration&lt;/a&gt; over
funding
&lt;a href="https://twitter.com/story645/status/1117567608222564353"&gt;maintenance and improvement&lt;/a&gt;
for open source tools
&lt;a href="https://twitter.com/jnuneziglesias/status/1131080022750519296"&gt;underpinning&lt;/a&gt;
an enormous amount of science. (Even a month later, reading Andreas's tweet
makes me want to scream in frustration in the crowded train in which I write
this. Not impactful enough!?) Given the buzz around CZI's call, I wonder
whether CZI has enough staff to even read through the avalanche of applications
it will receive. My hope is that national funding bodies will take notice and
start funding open source maintenance work, as Germany has
&lt;a href="https://www.dfg.de/en/research_funding/programmes/infrastructure/lis/funding_opportunities/call_proposal_software/"&gt;recently done&lt;/a&gt;.
(It's unclear to me whether Germany has repeated that call since 2016. If you
know, please leave a comment below!)&lt;/p&gt;
&lt;p&gt;Since I'm an academic, for a long time I was stuck in the mindset that granting
bodies need to step up and recognise that &lt;em&gt;established&lt;/em&gt; open source is a
valuable thing to fund, and this is certainly true. More recently though, as I
dipped my toes into some private sector consulting, it occurred to me that the
situation is completely ridiculous — “Hey look at this incredible thing we
built, it's serving thousands to millions of scientists, can we get some money
to keep it going?” “It's done and it's free, why would we give you money?” It's
a fundamental misunderstanding at the very top of funding agencies of how
software maintenance works.&lt;/p&gt;
&lt;p&gt;There are plenty of people that &lt;em&gt;do&lt;/em&gt; understand the value of the software,
though: its users, many of whom are in a position to use grant money to
help maintain it. But open source developers have not thus far made it
easy to &lt;del&gt;donate money to&lt;/del&gt; invest in their projects. I'm not talking about a
little donate button on the homepage, which will never substantially support
a project. Rather, open source projects should &lt;em&gt;sell&lt;/em&gt; a product, be it a
maintenance contract, a support contract, or a new feature development, or even
a &lt;em&gt;feel-good&lt;/em&gt; contract. But “buying” an open source software package should
feel exactly like buying closed-source software, and to University purchasing
departments it should look exactly like a normal software purchase.&lt;/p&gt;
&lt;p&gt;As I was thinking about this idea, I moved my blog from wordpress.com, for
which I was paying $100 per year, to &lt;a href="https://getnikola.com/"&gt;Nikola&lt;/a&gt;, an
open source static site generator written in Python. I &lt;em&gt;love&lt;/em&gt; Nikola, and was
ready to put my $100 towards the project, but to my surprise, even a donate
button was absent. As Andreas Mueller and Chris Holdgraf have
&lt;a href="https://twitter.com/choldgraf/status/1070672075692613636"&gt;pointed out&lt;/a&gt;,
many projects haven't even &lt;em&gt;considered&lt;/em&gt; what they would do with a large influx
of money if they got one. That needs to change.&lt;/p&gt;
&lt;p&gt;Thankfully, organizations including
&lt;a href="https://mail.python.org/pipermail/numpy-discussion/2019-April/079326.html"&gt;NumFOCUS&lt;/a&gt;,
&lt;a href="https://www.quansight.com/open-source-support"&gt;QuanSight&lt;/a&gt;,
&lt;a href="https://tidelift.com/"&gt;Tidelift&lt;/a&gt;, and others are working hard to rectify
this situation. I'm trying to help where I can and I really look forward to
seeing what they come up with.&lt;/p&gt;
&lt;p&gt;If you are interested in this topic, we are
&lt;a href="https://github.com/numfocus/scipy-2019-funding-foss-bof"&gt;organising&lt;/a&gt; a Birds
of a Feather session to take place at the
&lt;a href="https://www.scipy2019.scipy.org/"&gt;SciPy 2019&lt;/a&gt; in Austin in July.
Please join us!&lt;/p&gt;&lt;/div&gt;</description><category>open-source</category><category>Planet SciPy</category><category>programming</category><category>Python</category><category>science</category><guid>https://ilovesymposia.com/2019/05/28/why-citations-are-not-enough-for-open-source-software/</guid><pubDate>Tue, 28 May 2019 08:41:54 GMT</pubDate></item><item><title>Why you should cite open source tools</title><link>https://ilovesymposia.com/2019/05/02/why-you-should-cite-open-source-tools/</link><dc:creator>Juan Nunez-Iglesias</dc:creator><description>&lt;div&gt;&lt;p&gt;Every now and then, a moment or a sentence in a conversation sticks out at you,
and lodges itself in the back of your brain for months or even years. In this
case, the sentence is a tweet, and I fear that the only way to dislodge it is
to talk about it publicly.&lt;/p&gt;
&lt;p&gt;Last year, I complained on Twitter that a very prominent paper that was getting
lots of attention used scikit-image, but failed to cite our paper. (Or the
papers corresponding to many other open source packages.) I continued that
scientists developing open source software &lt;em&gt;depend&lt;/em&gt; on these citations to
continue their work. (More on this in another post...) One response was that
surely the developers of the open source scientific Python stack were not
scientists per se, and that citations were not a priority for them.&lt;/p&gt;
&lt;p&gt;I still sigh internally when I think of it.&lt;/p&gt;
&lt;p&gt;That tweet manifests a pervasive perception that open source scientific
software is written by God-like figures. These massively experienced software
developers have easy access to funds for their work, and are at the service of
all the other scientists, who are their users. I used to share this perception,
but it is utterly false.&lt;/p&gt;
&lt;!-- TEASER_END --&gt;

&lt;p&gt;I certainly hadn't thought about the funding question, but I did think of
packages like NumPy and SciPy as being written by "pros" (whatever that means),
whose main job was to produce these amazing libraries.&lt;/p&gt;
&lt;p&gt;My ideas started to change only after I attended the SciPy 2012 conference, and
I was invited to the scikit-image sprint by its lead author, &lt;a href="https://mentat.za.net/"&gt;Stéfan van der
Walt&lt;/a&gt;. I was totally starstruck, and even after meeting
him I just assumed his job was "open source guru", or somesuch. It is only
later that I learned that he was a postdoctoral researcher and lecturer in
applied mathematics at Stellenbosch University, in South Africa. As with other
academics, his main job was to produce research and to teach. Open source was
something he produced on the side.&lt;/p&gt;
&lt;p&gt;(Years later, as we were scrambling to finish the final chapter for
&lt;a href="http://elegant-scipy.org"&gt;Elegant SciPy&lt;/a&gt;, Stéfan revamped the build
infrastructure for our book — the scripts and configuration files that
converted the Markdown text we were writing to executed code and html. "You
have poor prioritisation skills, you know?", I taunted. He responded: "Yeah.
But a lot of the SciPy documentation toolchain exists because of that." And
yes: the revamped scripts ended up being extremely useful during the editing
phase with O'Reilly.)&lt;/p&gt;
&lt;p&gt;After the conference I continued to contribute to scikit-image, eventually
joining the core development team, but throughout the process I continued to
feel like an &lt;a href="https://en.wikipedia.org/wiki/Impostor_syndrome"&gt;impostor&lt;/a&gt;,
someone who had somehow managed to gain entrance to this hallowed and
otherworldly community despite his inferior skills and knowledge.&lt;/p&gt;
&lt;p&gt;Only after years of interacting with this community did I internalise the fact
that nearly all of this software stack has been produced by practising
scientists who took the extra care and effort to ensure that their code was
robust, well-tested, and easily accessible to all. Despite a recent influx of
interest and contributions from industry, as far as I can tell, most
contributions to the SciPy stack still come from practising scientists in
academia.&lt;/p&gt;
&lt;p&gt;I know this now, but until recently I was suffering from another fallacy: that
if I knew it, clueless as I was, then surely everybody knew it. That tweet
disabused me of that notion, and this is why you're reading this. I hope it is
instructive, and that you'll find it worth sharing widely.&lt;/p&gt;
&lt;p&gt;This idea, that the SciPy stack is made by active scientists, is important
because it affects how this work can be supported. Sadly, neither university
hierarchies nor national funding bodies recognise code as valuable output.
(There are some
&lt;a href="https://twitter.com/ethanwhite/status/1083008449242378240"&gt;exceptions&lt;/a&gt;, but
this remains the norm.) By and large, the only things that count are papers,
grants, and, to a lesser extent, teaching evaluation scores. &lt;/p&gt;
&lt;p&gt;So, yes, many of the original contributors to the SciPy libraries are now in
industry. But they were academics at the time that they contributed and were
driven out because academia did not value their contributions. Academics should
not have to sacrifice their careers to contribute to open source. Citations to
software papers are an imperfect solution to this problem (again, more soon),
but they sure as hell are better than nothing.&lt;/p&gt;
&lt;p&gt;So, if you are a user of open source tools in the Scientific Python stack, I
have two requests for you:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;When you publish your work, cite every library that you import. Most
   scientific software has a notice on their homepage or README file pointing
   to a paper you can cite. By definition, if you've imported a library, you've
   found it useful, and if you've found it useful, then you probably care about
   supporting its authors. This is a small way you can contribute to their
   success.&lt;/li&gt;
&lt;li&gt;You are good enough to contribute. If you have an issue
   with an open source package you are using, look at the source code. Submit
   an issue to the project's bug tracker (usually GitHub). And try your hand at
   fixing it. The software's authors will usually offer guidance on
   how to do this, and you will improve your own skills as a result. Good
   software development practices is one of the most transferrable skills you
   can gain.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Of course, citations alone will not solve the wider problem that open source
software is chronically undervalued. I have many thoughts about how open source
should be supported, especially in science, but I'll expand on that in an
upcoming post.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Update:&lt;/strong&gt; I'm adding a link to a related post: &lt;a href="http://ilovesymposia.com/2018/06/20/what-do-scientists-know-about-open-source/"&gt;What do scientists know about
open source?&lt;/a&gt;&lt;/p&gt;&lt;/div&gt;</description><category>open-source</category><category>Planet SciPy</category><category>programming</category><category>Python</category><category>science</category><guid>https://ilovesymposia.com/2019/05/02/why-you-should-cite-open-source-tools/</guid><pubDate>Thu, 02 May 2019 02:31:30 GMT</pubDate></item><item><title>Summer school announcement: 2nd Advanced Scientific Programming in Python (ASPP) Asia Pacific!</title><link>https://ilovesymposia.com/2018/08/30/summer-school-announcement-2nd-advanced-scientific-programming-in-python-aspp-asia-pacific/</link><dc:creator>Juan Nunez-Iglesias</dc:creator><description>&lt;div&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;The Advanced Scientific Programming in Python (ASPP) summer school has had &lt;a href="https://scipy-school.org/archives"&gt;10 successful iterations&lt;/a&gt; in Europe and &lt;a href="https://python.g-node.org/aspp-asia-pacific-2018/"&gt;one iteration here in Melbourne&lt;/a&gt; earlier this year. Another European iteration is starting next week in Camerino, Italy.&lt;/p&gt;
&lt;p&gt;Now, thanks to the generous sponsorship of &lt;a href="https://www.csiro.au"&gt;CSIRO&lt;/a&gt;, and the efforts of &lt;a href="http://biology.anu.edu.au/people/benjamin-schwessinger"&gt;Benjamin Schwessinger&lt;/a&gt; and &lt;a href="https://twitter.com/DataNerdery"&gt;Genevieve Buckley&lt;/a&gt;, two alumni from the Melbourne school, and &lt;a href="https://people.csiro.au/M/K/Kerensa-Mcelroy"&gt;Kerensa McElroy&lt;/a&gt;, Agriculture Data School Coordinator at CSIRO, the Asia Pacific fork of ASPP gets its second iteration in &lt;strong&gt;Canberra&lt;/strong&gt;, &lt;strong&gt;Jan 20-27, 2019&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;/p&gt;&lt;h3&gt;Key details&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;The workshop runs &lt;strong&gt;January 20-27, 2019&lt;/strong&gt; at the &lt;strong&gt;Australian National University&lt;/strong&gt; in Canberra, Australia.&lt;/li&gt;
&lt;li&gt;topics include git, contributing to open source software with github, testing, debugging, profiling, advanced NumPy, Cython, and data visualisation.&lt;/li&gt;
&lt;li&gt;hands-on learning using pair programming&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;free to attend&lt;/strong&gt; (but students are responsible for travel, accommodation, and meals)&lt;/li&gt;
&lt;li&gt;30 student places, to be selected competitively&lt;/li&gt;
&lt;li&gt;application deadline is &lt;strong&gt;Oct 7, 2018&lt;/strong&gt;, 23:59 &lt;a href="https://www.timeanddate.com/time/zones/aoe"&gt;Anywhere On Earth&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;website: https://scipy-school.org&lt;/li&gt;
&lt;li&gt;FAQ: https://scipy-school.org/faq&lt;/li&gt;
&lt;li&gt;apply: https://scipy-school.org/applications&lt;/li&gt;
&lt;/ul&gt;

&lt;!-- TEASER_END --&gt;

&lt;h3&gt;Background&lt;/h3&gt;

&lt;p&gt;Three years ago, I had the privilege of teaching the 2015 ASPP school in Munich. It turned out to be a fantastic teaching experience (I have taught in 2 more since), and more importantly, it was a fantastic experience for the students. Students are selected for the school to fit a certain profile, neither too novice nor too advanced. As such, participants selected for the school are almost guaranteed to learn a great deal.&lt;/p&gt;
&lt;p&gt;Indeed, almost every iteration of the school has been co-organised by former students. Sure enough, with the help of two students from the Melbourne instance, we will be able to have a new iteration in Canberra this January.&lt;/p&gt;
&lt;h3&gt;Course description&lt;/h3&gt;

&lt;p&gt;Scientists spend increasingly more time writing, maintaining, and debugging software. While techniques for doing this efficiently have evolved, only few scientists have been trained to use them. As a result, instead of doing their research, they spend far too much time writing deficient code and reinventing the wheel. In this course we will present a selection of advanced programming techniques and best practices that are standard in industry, but especially tailored to the needs of a programming scientist. Lectures are devised to be interactive and to give the students enough time to acquire direct hands-on experience with the materials. Students will work in pairs throughout the school and will team up to practice the newly learned skills in a real programming project — an entertaining computer game.&lt;/p&gt;
&lt;p&gt;We use the Python programming language for the entire course. Python works as a simple programming language for beginners, but more importantly, it also works great in scientific simulations and data analysis. We show how clean language design, ease of extensibility, and the great wealth of open source libraries for scientific computing and data visualization are driving Python to becoming a standard tool for scientists.&lt;/p&gt;
&lt;h3&gt;Who is eligible?&lt;/h3&gt;

&lt;p&gt;This school is targeted at Master/PhD students, postdocs, and academic staff and technicians from all areas of science. Competence in Python or in another language such as Java, C/C++, MATLAB, or Mathematica is absolutely required. Basic knowledge of Python and of a version control system such as git, subversion, mercurial, or bazaar is assumed. Participants without any prior experience with Python and/or git should work through the proposed introductory material before the course.&lt;/p&gt;
&lt;p&gt;We have strived to get a pool of students that is international and gender-balanced, and have succeeded, with gender parity in the last five schools.&lt;/p&gt;
&lt;h3&gt;More questions&lt;/h3&gt;

&lt;p&gt;If you have any questions, contact &lt;a href="mailto:info@scipy-school.org"&gt;info@scipy-school.org&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Please circulate this announcement widely! And follow &lt;a href="https://twitter.com/scipyschool"&gt;@scipyschool&lt;/a&gt; for further developments.&lt;/p&gt;
&lt;p&gt;Juan.&lt;/p&gt;&lt;/div&gt;</description><category>conference</category><category>open-source</category><category>Planet SciPy</category><category>programming</category><category>Python</category><category>science</category><guid>https://ilovesymposia.com/2018/08/30/summer-school-announcement-2nd-advanced-scientific-programming-in-python-aspp-asia-pacific/</guid><pubDate>Thu, 30 Aug 2018 04:48:05 GMT</pubDate></item><item><title>What do scientists know about open source?</title><link>https://ilovesymposia.com/2018/06/20/what-do-scientists-know-about-open-source/</link><dc:creator>Juan Nunez-Iglesias</dc:creator><description>&lt;div&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;A friend recently pointed out this great talk by Matt Bernius, &lt;a href="https://community.redhat.com/blog/2018/05/what-college-students-know/"&gt;What students know and don't know about open source&lt;/a&gt;. If you have even a minor interest in open source it's worth a watch, but the gist is: in the US alone, there are about 200,000 students enrolled in a computer science major. Open source communities are a great space to learn real-world programming, so why don't these numbers translate into massive contributions to open source?&lt;/p&gt;
&lt;p&gt;At the core of the issue, Matt identifies two main problems: (1) colleges and universities simply don't teach open source, or even collaborative coding; and (2), many open source communities make newcomers feel unwelcome in a variety of ways.&lt;/p&gt;
&lt;p&gt;I want to comment about this in the context of programming in science. That is, programming where the code is not the main product, but rather a useful tool to obtain a scientific result, for example in biology or physics. Here, we still see relatively little contribution to open source, for related but different cultural issues.&lt;/p&gt;
&lt;!-- TEASER_END --&gt;

&lt;p&gt;I've sent my &lt;a href="https://ilovesymposia.com/2015/12/26/why-scientists-should-code-in-the-open/"&gt;scientists should code in the open&lt;/a&gt; post to a few people and the response from most remains sceptical. I hope this post will address some of their concerns.&lt;/p&gt;
&lt;p&gt;&lt;/p&gt;&lt;h3&gt;Scientific culture is ridiculously secretive&lt;/h3&gt;
&lt;p&gt;The most common objection is to my assertion that people won't scoop you by looking at your code. I remember a tweet (that I sadly can't find now) that really got to the gist of the problem. It went something like this:&lt;/p&gt;
&lt;blockquote&gt;
  Someone in science having a new idea: "Ooh, I hope I don't get scooped!"&lt;br&gt;
  Someone in open source having a new idea: "Ooh, I hope someone has implemented this already!"
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Update:&lt;/strong&gt; I found the source! It's &lt;a href="https://twitter.com/tweetotaler/status/884412302098915329"&gt;this tweet&lt;/a&gt; by Elizabeth Seiver.&lt;/p&gt;
&lt;p&gt;This is a huge gap in culture that won't soon go away, but there are encouraging steps towards narrowing it. For example, PLOS Biology, a leading journal, recently &lt;a href="http://twitter.com/PLOSBiology/status/958346565868978176"&gt;announced&lt;/a&gt; that they would consider "scooped" studies for publication within six months of the "scooping". That goes some way towards re-aligning incentives towards collaborative and open science.&lt;/p&gt;
&lt;p&gt;I've come across many collaborations that have started because of open source. I have not heard of someone getting scooped because of open source, but of course that sort of information would be hard to trace and come by. Several people did write to me that they were concerned about very specific groups rifling through their code expressly for the purpose of scooping them. For me it's hard to imagine someone even having that attitude, and my advice is that if you do face such a toxic community, it might be wise to change your chosen field of study.&lt;/p&gt;
&lt;p&gt;Nevertheless, I want to emphasise here that open source programming can take many forms, with the zip file attached to the paper being the lowest, coding in the open being the highest, and several other models in between. Any steps you can take towards the higher models will ultimately help you. My preferred mode for code that really does have to be private is to use a private GitHub repository, and &lt;em&gt;just make that repo public once the paper is accepted.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;A lot of people prefer the "code dump with no revision history" model of post-publication sharing, but this tosses out a lot of valuable information for people coming after you: what have you tried that didn't work? What issues did you have with the code? Have you considered coding in one style or another? The code dump model also makes you less likely to use GitHub in the first place, depriving you of an opportunity to learn some valuable real-world skills.&lt;/p&gt;
&lt;h3&gt;For coding, scientists have even more severe impostor syndrome&lt;/h3&gt;

&lt;p&gt;As I mentioned in my original post, and this I find completely uncontestable, &lt;em&gt;publishing shitty code is not a bad thing.&lt;/em&gt; &lt;em&gt;Everybody writes bad code,&lt;/em&gt; and nearly everybody knows it. Here's Hadley Wickham, creator of &lt;a href="https://dplyr.tidyverse.org"&gt;dplyr&lt;/a&gt;, &lt;a href="https://tidyr.tidyverse.org"&gt;tidy data&lt;/a&gt;, &lt;a href="https://ggplot2.tidyverse.org"&gt;ggplot2&lt;/a&gt;, among other things; in other words, someone who knows a thing or two about elegant code and about as close as one gets to coding royalty in science:&lt;/p&gt;
&lt;blockquote&gt;
  The only way to write great code is to write lots of shitty code first.
&lt;/blockquote&gt;

&lt;p&gt;Publishing your raw code is a good thing and will absolutely not be a black mark on your career. Indeed, in open source circles, it is often a bare GitHub contribution history that is a black mark. (And this is another problem, but in my opinion a better one.)&lt;/p&gt;
&lt;h3&gt;Scientists don't know about open source&lt;/h3&gt;

&lt;p&gt;If knowledge of open source is lacking in computer science, what chance does it have in other fields? The truth is that outreach and education need to become a massive part of open source culture, &lt;em&gt;especially&lt;/em&gt; in science.&lt;/p&gt;
&lt;p&gt;I credit &lt;a href="https://bids.berkeley.edu/people/st%C3%A9fan-van-der-walt"&gt;Stéfan van der Walt&lt;/a&gt; for my life in open source. After I gave a talk at SciPy 2012, he invited me to join the scikit-image sprint at the end of the conference. If it hadn't been for that, I probably would have just wandered around the hall, too shy to join any sprint (see "impostor syndrome", above), and my life would be very different right now.&lt;/p&gt;
&lt;p&gt;Anyway, at that point I'd made my code "open source", which meant it was on GitHub. I had only added a license to submit to the conference. As a reminder, unlicensed code &lt;a href="http://www.astrobetter.com/blog/2014/03/10/the-whys-and-hows-of-licensing-scientific-code/"&gt;doesn't count as open source&lt;/a&gt;. But I had never really collaborated in open source. My idea of collaboration was my workflow with my colleague: a single branch (master), from which we both pulled and to which we both pushed. When I sat down with Stéfan and &lt;a href="http://tonysyu.github.io"&gt;Tony Yu&lt;/a&gt;, and I figured what I wanted to work on, I asked: "So, should I just push to master, or what?" I still remember, with some embarrassment, the dubious look Stéfan and Tony exchanged, as they silently figured out which of them would introduce this newbie to &lt;a href="https://help.github.com/articles/creating-a-pull-request/"&gt;pull requests&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;But that's the thing: I shouldn't feel embarrassment. Scientists for the most part don't get introduced to coding in their education, much less to open source.&lt;/p&gt;
&lt;h3&gt;What can scientists in open source do?&lt;/h3&gt;

&lt;p&gt;A lesson from my continued contributions to the SciPy ecosystem, I hope, is that some light mentorship can yield enormous dividends later on. Stéfan and Tony took the time to walk me through the open source contribution process, when they could have dismissively sent me a link to some page explaining it. I'm a big fan of writing good documents for newcomers, but nothing beats a good hand-holding. It's very easy for me to imagine an alternate reality where I had not felt welcome or rewarded by the scikit-image project and my life had not taken this productive turn.&lt;/p&gt;
&lt;p&gt;Continuing on imaginary themes, it is only slightly less plausible that the open source scientific world should be awash with new contributors at every level of science. How do we turn this dream into a reality?&lt;/p&gt;
&lt;p&gt;If you are a scientist and this post is among your first encounters with the term "open source", and you think you might be interested in learning more, here are a few things I recommend, in order of easiest to hardest:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Read the &lt;a href="https://github.com/elegant-scipy/elegant-scipy/blob/master/markdown/preface.markdown"&gt;preface&lt;/a&gt; and &lt;a href="https://github.com/elegant-scipy/elegant-scipy/blob/master/markdown/epilogue.markdown"&gt;epilogue&lt;/a&gt; of my book with Stéfan and &lt;a href="http://harrietdashnow.com"&gt;Harriet Dashnow&lt;/a&gt;. (Free online!) I feel a bit icky recommending my own book, but why repeat myself? In those chapters I tried to distill my thoughts on joining the SciPy community, which is a fantastic, rewarding space in which to do open source programming as a scientist. I expect many things we wrote generalise well to e.g. &lt;a href="https://www.tidyverse.org"&gt;the tidyverse&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Look for upcoming &lt;a href="https://software-carpentry.org"&gt;software carpentry&lt;/a&gt; workshops near you. These are free two-day programming boot camps to introduce you to computational thinking, and, crucially, to version control with git.&lt;/li&gt;
&lt;li&gt;Go to a &lt;a href="https://conference.scipy.org"&gt;SciPy conference&lt;/a&gt;. I know of SciPy, EuroSciPy, and SciPy India, but I have a vague memory of offshoots in Africa and South America.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are in a boat similar to mine (intermediate/advanced open source contributor in science), and you feel like you would like your work to feel a bit more crowded, I can tell you what I'm going to be doing in response to this talk:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sign up to deliver (more) software carpentry training (or similar). Getting the word out is the number one thing.&lt;/li&gt;
&lt;li&gt;In software carpentry, emphasise the role of git in collaboration. (I think the official program does not go far enough in this direction, and focuses instead on the initial linear history.)&lt;/li&gt;
&lt;li&gt;If you are located in a university, talk to your CS department to see whether they have any courses in open source development. If not, see whether you can guest lecture in a suitable course to make students aware of the open source opportunities out there.&lt;/li&gt;
&lt;li&gt;Similarly, follow up software carpentry with more advanced sessions on open source collaboration. I gained an enormous fraction of my programming skills from collaborating on open source. I really think there is no better tool for long-term learning in this space. An idea that I'd like to try out is to curate a bunch of open issues on prominent repos and get SWC students to sprint on them for a day&lt;sup id="fnref-1111-1"&gt;&lt;a href="https://ilovesymposia.com/2018/06/20/what-do-scientists-know-about-open-source/#fn-1111-1" class="jetpack-footnote"&gt;1&lt;/a&gt;&lt;/sup&gt;. I know about the "good first issue" tag on GitHub. Unfortunately, my experience with it is mixed. I think many repos are overly optimistic with theirs (this includes scikit-image), and, furthermore, a large proportion of these tagged issues get "claimed" quickly — and often half-heartedly!&lt;/li&gt;
&lt;li&gt;Write, write, write! Did you get a cool PR merged? Write a blog post about it! Or at least tweet! We need to get the message out that writing PRs is for everyone. =)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have any further ideas, I'd love to hear them.&lt;/p&gt;
&lt;div class="footnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;

&lt;li id="fn-1111-1"&gt;
Actually I drafted this post a while back, and tried this yesterday, with mixed success. I'll write about that experience soon. ;) &lt;a href="https://ilovesymposia.com/2018/06/20/what-do-scientists-know-about-open-source/#fnref-1111-1"&gt;↩&lt;/a&gt;
&lt;/li&gt;

&lt;/ol&gt;
&lt;/div&gt;

&lt;p&gt;&lt;/p&gt;&lt;/div&gt;</description><category>open-source</category><category>Planet SciPy</category><category>programming</category><category>science</category><guid>https://ilovesymposia.com/2018/06/20/what-do-scientists-know-about-open-source/</guid><pubDate>Wed, 20 Jun 2018 11:19:46 GMT</pubDate></item><item><title>1st ASPP Asia Pacific evaluation survey</title><link>https://ilovesymposia.com/2018/04/09/1st-aspp-asia-pacific-evaluation-survey/</link><dc:creator>Juan Nunez-Iglesias</dc:creator><description>&lt;div&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;In January of 2018, we had the first &lt;a href="http://python.g-node.org"&gt;ASPP summer school&lt;/a&gt; outside of Europe. (This was a parallel workshop to the European one, which will be held in Italy in September 2018.) In general, it was a great success, with some caveats that we will elaborate on below.&lt;/p&gt;
&lt;p&gt;First we want to note that this school was a bit different than the European ones, in that we only had attendees from Australian institutions, where the European school has broad international representation, including some from out of Europe. This was in some ways inevitable, as it is more expensive to travel to Australia from almost anywhere than to travel within Europe. On the other hand, we advertised relatively late, and we were unable to secure travel grants during the advertising period, so there is hope that a future edition would be able to attract a more international crowd from the Asia Pacific region.&lt;/p&gt;
&lt;!-- TEASER_END --&gt;

&lt;p&gt;Given all this, there was a question as to whether we would be able to capture the atmosphere of the school, which normally sees the students living together and socialising for basically the whole week. In this case, most students just went home after classes were finished. But although some of that atmosphere was missing, by the end of the week we did manage to get some close links between all the students and the faculty. The evaluations below show that most of the value of the school was preserved.&lt;/p&gt;
&lt;p&gt;We note that 100% of the respondents (29/30 of the students) would recommend the course to their peers. So, although some lectures were better received than others, and although the programming project was not universally loved, we managed to provide value for everyone. All of this is in line with the evaluations at previous schools (available at https://python.g-node.org/wiki/archives.html).&lt;/p&gt;
&lt;p&gt;The project, which consists of programming a videogame bot, is controversial every year, but, consistently, more people like it than don't, and people get to practice git, pair programming, and programming as a team, which is the single most difficult skill to practice when programming for science. Indeed when we walk around during the project programming sessions, we see people extremely engaged in what they are coding. It's difficult to imagine a scientific problem engaging such diverse people as the school's attendees (which come from very disparate scientific fields).&lt;/p&gt;
&lt;p&gt;Of all the feedback, two particular statements, we hope from people in the same project group, broke our hearts. We decided not to include them in this report, because they might be easy to de-anonymise by group members, but they boil down to the following: a group member, by being combative and rude to others in their team, and deciding to essentially complete the project by themselves, ruined the programming project for all of their team members, with some even feeling that they were not good enough to contribute. This is tragic, because we want everyone in the school to feel &lt;em&gt;empowered&lt;/em&gt; to do anything at all in Python.&lt;/p&gt;
&lt;p&gt;Absolutely every student has something to offer in this project. Here, as in life, teams are comprised of members of varying skills. But we know from our selection that everyone has the skills to contribute (and this is confirmed by the fact that most attendees, for most lectures, felt that the difficulty level was "just right"). So if a student felt inadequate, it can only be because of the toxic team member.&lt;/p&gt;
&lt;p&gt;Ned Batchelder recently wrote an excellent &lt;a href="https://nedbatchelder.com/blog/201711/toxic_experts.html"&gt;blog post&lt;/a&gt; about what he calls "Toxic experts" and what Tiziano Zito calls, somewhat more bluntly, "Arrogant assholes". (In discussions about this post, Tiziano and others noted that one does not have to be an expert to be toxic, or arrogant, or an asshole. No matter: the points below apply equally to anyone meeting any of the above characteristics &lt;em&gt;regardless&lt;/em&gt; of expertise.)&lt;/p&gt;
&lt;p&gt;The feedback we received should serve as a warning to selection committees and hiring managers everywhere about how damaging it is to allow such a person into your ranks. Due to the anonymous nature of the survey, we can't tell whether there was one or two toxic experts in our midst, but if it's one, they soured the school for five other people. If it's two, then that's ten people, a third of the school, that might have had a terrible experience. The problem with toxic experts is that they can so quickly cause damage to so many others. Thus, even if they are a mythical "10x engineer", &lt;strong&gt;they are not worth it.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Literally nothing that the above-described team member could have done, coding-wise, could make up for the damage they caused. Despite their strong opinions, they missed the entire point of the programming project, which is not to win a medal, but to &lt;em&gt;learn about working in a team.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;We try to avoid toxic experts in our selection process for the school, but they slip through every so often. In response to this feedback, we will aim to be even more vigilant in our selection, and also make the aims of the project &lt;em&gt;as a learning exercise&lt;/em&gt; more explicit during its introduction. We will also make sure to be more aware of group interactions during the actual school; we apologise to the students involved that we did not catch this behaviour this time. We are truly sorry.&lt;/p&gt;
&lt;p&gt;If you are in the position of being an expert during a school or workshop, don't go it alone. That is a waste of your time, because you can do a programming project on your own whenever you damn well please. Slow down, and think instead about practicing your teaching and mentoring skills. They are also important in life, and, in many contexts, they are your responsibility.&lt;/p&gt;
&lt;p&gt;You can access the full survey results &lt;a href="https://python.g-node.org/wiki/_media/evaluation_survey_2018_melbourne.pdf"&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;-- Juan, and the Organisers.&lt;/p&gt;&lt;/div&gt;</description><category>conference</category><category>programming</category><category>Python</category><category>science</category><guid>https://ilovesymposia.com/2018/04/09/1st-aspp-asia-pacific-evaluation-survey/</guid><pubDate>Mon, 09 Apr 2018 03:29:53 GMT</pubDate></item><item><title>SciPy's new LowLevelCallable is a game-changer</title><link>https://ilovesymposia.com/2017/03/12/scipys-new-lowlevelcallable-is-a-game-changer/</link><dc:creator>Juan Nunez-Iglesias</dc:creator><description>&lt;div&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;... and combines rather well with that other game-changing library I like, &lt;a href="https://ilovesymposia.com/2016/12/20/numba-in-the-real-world/"&gt;Numba&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I've &lt;a href="https://ilovesymposia.com/2015/12/10/the-cost-of-a-python-function-call/"&gt;lamented before&lt;/a&gt; that function calls are expensive in Python, and that this severely hampers many functions that &lt;em&gt;should&lt;/em&gt; be insanely useful, such as SciPy's &lt;a href="https://docs.scipy.org/doc/scipy/reference/generated/scipy.ndimage.generic_filter.html#scipy.ndimage.generic_filter"&gt;&lt;code&gt;ndimage.generic_filter&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;!-- TEASER_END --&gt;

&lt;p&gt;To illustrate this, let's look at image &lt;em&gt;erosion&lt;/em&gt;, which is the replacement of each pixel in an image by the minimum of its neighbourhood. &lt;code&gt;ndimage&lt;/code&gt; has a fast C implementation, which serves as a perfect benchmark against the generic version, using a generic filter with &lt;code&gt;min&lt;/code&gt; as the operator. Let's start with a 2048 x 2048 random image:&lt;/p&gt;
&lt;pre class="code literal-block"&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;numpy&lt;/span&gt; &lt;span class="kn"&gt;as&lt;/span&gt; &lt;span class="nn"&gt;np&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;image&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/pre&gt;


&lt;p&gt;and a neighbourhood “footprint” that picks out the pixels to the left and right, and above and below, the centre pixel:&lt;/p&gt;
&lt;pre class="code literal-block"&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;footprint&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;array&lt;/span&gt;&lt;span class="p"&gt;([[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;                       &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;                       &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt; &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;


&lt;p&gt;Now, we measure the speed of &lt;code&gt;grey_erosion&lt;/code&gt; and &lt;code&gt;generic_filter&lt;/code&gt;. Spoiler alert: it’s not pretty.&lt;/p&gt;
&lt;pre class="code literal-block"&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;scipy&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ndimage&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;ndi&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;timeit&lt;/span&gt; &lt;span class="n"&gt;ndi&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grey_erosion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="n"&gt;loops&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;118&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;timeit&lt;/span&gt; &lt;span class="n"&gt;ndi&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;generic_filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;min&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;27&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;
&lt;/pre&gt;


&lt;p&gt;As you can see, with Python functions, &lt;code&gt;generic_filter&lt;/code&gt; is unusable for anything but the tiniest of images.&lt;/p&gt;
&lt;p&gt;A few months ago, I was
&lt;a href="https://groups.google.com/a/continuum.io/d/msg/numba-users/HMg_65R8KZE/RnysYokGAwAJ"&gt;trying&lt;/a&gt;
to get around this by using Numba-compiled functions, but the way to feed C
functions to SciPy was different depending on which part of the library you
were using. &lt;code&gt;scipy.integrate&lt;/code&gt; used ctypes, while &lt;code&gt;scipy.ndimage&lt;/code&gt; used &lt;code&gt;PyCObjects&lt;/code&gt; or
&lt;code&gt;PyCapsules&lt;/code&gt;, depending on your Python version, and Numba only supported the
former method at the time. (Plus, this topic starts to stretch my understanding
of low-level Python, so I felt there wasn’t much I could do about it.)&lt;/p&gt;
&lt;p&gt;Enter &lt;a href="https://github.com/scipy/scipy/pull/6509"&gt;this pull request&lt;/a&gt; to SciPy
from Pauli Virtanen, which is live in the most recent SciPy version, 0.19. It
unifies all C-function interfaces within SciPy, and Numba already supports this
format. It takes a bit of gymnastics, but it works! It really works!&lt;/p&gt;
&lt;p&gt;(By the way, the release is full of little gold nuggets. If you use SciPy at
all, the release notes are well worth a read.)&lt;/p&gt;
&lt;p&gt;First, we need to define a C function of the appropriate signature. Now, you
might think this is the same as the Python signature, taking in an array of
values and returning a single value, but that would be too easy! Instead, we
have to go back to some C-style programming with pointers and array sizes. From
the &lt;code&gt;generic_filter&lt;/code&gt; documentation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This function also accepts low-level callback functions with one of the
following signatures and wrapped in scipy.LowLevelCallable:&lt;/p&gt;
&lt;p&gt;&lt;code&gt;c
int callback(double *buffer, npy_intp filter_size, 
             double *return_value, void *user_data)
int callback(double *buffer, intptr_t filter_size, 
             double *return_value, void *user_data)&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;The calling function iterates over the elements of the input and output
arrays, calling the callback function at each element. The elements within
the footprint of the filter at the current element are passed through the
buffer parameter, and the number of elements within the footprint through
filter_size. The calculated value is returned in return_value. user_data is
the data pointer provided to scipy.LowLevelCallable as-is.&lt;/p&gt;
&lt;p&gt;The callback function must return an integer error status that is zero if
something went wrong and one otherwise. &lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(Let’s leave aside that crazy reversal of Unix convention of the past 50 years
in the last paragraph, except to note that our function must return 1 or it
will be killed.)&lt;/p&gt;
&lt;p&gt;So, we need a Numba cfunc that takes in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a double pointer pointing to the values within the footprint,&lt;/li&gt;
&lt;li&gt;a pointer-sized integer that specifies the number of values in the footprint,&lt;/li&gt;
&lt;li&gt;a double pointer for the result, and&lt;/li&gt;
&lt;li&gt;a void pointer, which could point to additional parameters, but which we can ignore for now.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The Numba type names are listed in &lt;a href="http://numba.pydata.org/numba-doc/dev/reference/types.html#numba-types"&gt;this page&lt;/a&gt;. Unfortunately, at the time of
writing, there’s no mention of how to make pointers there, but finding such a
reference &lt;a href="http://numba.pydata.org/numba-doc/dev/user/cfunc.html#signature-specification"&gt;was not too hard&lt;/a&gt;. (Incidentally, it would make a good contribution to
Numba’s documentation to add CPointer to the Numba types page.)&lt;/p&gt;
&lt;p&gt;So, armed with all that documentation, and after much trial and error, I was
finally ready to write that C callable:&lt;/p&gt;
&lt;pre class="code literal-block"&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;numba&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cfunc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;carray&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;numba.types&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;intc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;intp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;voidptr&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;numba.types&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;CPointer&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; 
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; 
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nd"&gt;@cfunc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;CPointer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;intp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;             &lt;span class="n"&gt;CPointer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;voidptr&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;nbmin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;values_ptr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;len_values&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;     &lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;carray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;values_ptr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;len_values&lt;/span&gt;&lt;span class="p"&gt;,),&lt;/span&gt; &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;     &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;inf&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;     &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;         &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;             &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;
&lt;span class="o"&gt;...&lt;/span&gt;     &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/pre&gt;


&lt;p&gt;The only other tricky bits I had to watch out for while writing that function were as follows:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;remembering that there’s two ways to de-reference a pointer in C: &lt;code&gt;*ptr&lt;/code&gt;,
   which is not valid Python and thus not valid Numba, and &lt;code&gt;ptr[0]&lt;/code&gt;. So, to place
   the result at the given double pointer, we use the latter syntax. (If you
   prefer to use Cython, the same rule applies.)&lt;/li&gt;
&lt;li&gt;Creating an array out of the &lt;code&gt;values_ptr&lt;/code&gt; and &lt;code&gt;len_values&lt;/code&gt; variables, as shown
   here. That’s what enables the &lt;code&gt;for v in values&lt;/code&gt; Python-style access to the
   array.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Ok, so now what you’ve been waiting for. How did we do? First, to recap, the original benchmarks:&lt;/p&gt;
&lt;pre class="code literal-block"&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;scipy&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ndimage&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;ndi&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;timeit&lt;/span&gt; &lt;span class="n"&gt;ndi&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grey_erosion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="n"&gt;loops&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;118&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;timeit&lt;/span&gt; &lt;span class="n"&gt;ndi&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;generic_filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;min&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;27&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;
&lt;/pre&gt;


&lt;p&gt;And now, with our new Numba cfunc:&lt;/p&gt;
&lt;pre class="code literal-block"&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;timeit&lt;/span&gt; &lt;span class="n"&gt;ndi&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;generic_filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;LowLevelCallable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nbmin&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ctypes&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;footprint&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="n"&gt;loops&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;113&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt;
&lt;/pre&gt;


&lt;p&gt;That's right: it's even marginally &lt;em&gt;faster&lt;/em&gt; than the pure C version! I almost cried when I ran that.&lt;/p&gt;


&lt;hr&gt;

&lt;p&gt;Higher-order functions, ie functions that take other functions as input, enable powerful, concise, elegant &lt;a href="https://ilovesymposia.com/2014/06/24/a-clever-use-of-scipys-ndimage-generic_filter-for-n-dimensional-image-processing/"&gt;expressions&lt;/a&gt; of various algorithms. Unfortunately, these have been hampered in Python for large-scale data processing because of Python's function call overhead. SciPy's latest update goes a long way towards redressing this.&lt;/p&gt;&lt;/div&gt;</description><category>Numba</category><category>open-source</category><category>Planet SciPy</category><category>programming</category><category>Python</category><category>science</category><category>SciPy</category><guid>https://ilovesymposia.com/2017/03/12/scipys-new-lowlevelcallable-is-a-game-changer/</guid><pubDate>Sun, 12 Mar 2017 03:41:41 GMT</pubDate></item><item><title>Brian Greene on the Colbert Report</title><link>https://ilovesymposia.com/2008/05/29/brian-greene-on-the-colbert-report/</link><dc:creator>Juan Nunez-Iglesias</dc:creator><description>&lt;div&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;I promise sometime soon I'll write something &lt;em&gt;not&lt;/em&gt; about someone else's videos! But for now, enjoy theoretical physicist &lt;a href="http://www.comedycentral.com/colbertreport/videos.jhtml?videoId=167386"&gt;Brian Greene on the Colbert Report&lt;/a&gt;. Stephen drives an excellent interview, as usual, and proves &lt;a href="http://www.comedycentral.com/colbertreport/videos.jhtml?videoId=76296"&gt;yet again&lt;/a&gt; that he either knows a good deal of science, or he does his homework before talking about it. As a result, science coverage on the Colbert Report is invariably excellent. &lt;a href="http://ilovesymposia.wordpress.com/wp-admin/post.php?action=edit&amp;amp;post=4&amp;amp;message=4"&gt;&lt;/a&gt;&lt;/p&gt;&lt;/div&gt;</description><category>brian greene</category><category>physics</category><category>science</category><category>stephen colbert</category><category>Video</category><guid>https://ilovesymposia.com/2008/05/29/brian-greene-on-the-colbert-report/</guid><pubDate>Thu, 29 May 2008 07:21:18 GMT</pubDate></item></channel></rss>