<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IWSG</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>nanoHUB user behavior: moving from retrospective statistics to actionable behavior analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gerhard Klimeck</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gustavo A. Valencia-Zapata</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nathan Denny</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lynn K. Zentner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael G. Zentner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Network for Computational Nanotechnology</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rosen Center for Advanced Computing</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Purdue University</institution>
          ,
          <addr-line>West Lafayette, IN 47907</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>12</volume>
      <fpage>12</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>nanoHUB annually serves 17,000+ registered users with over 1 million simulations. In the past, we have used data analytics to demonstrate that nanoHUB can be a powerful scientific knowledge sharing platform. We used retrospective data analytics to show how simulation tools were used in structured education and how simulation tools were used in novel research. With the use of such retrospective analytics, we have made strategic decisions in terms of tool and content developments and justified continued nanoHUB investments by the US National Science Foundation (NSF). As we migrate towards a sustainable nanoHUB we must embrace similar processes pursued by in similar platforms such as Uber or AirBnB: we need to create actionable data analytics that can rapidly support user experience and help grow the supply in the two-sided market platform - we need to improve the experience of providers as well as end-users. This paper describes some aspects on how we pursue user behavior analysis inside the virtual worlds of nanotechnology simulation tools. From such user behavior we plan to derive actionable analytics that influence user behaviors as they interact with nanoHUB.</p>
      </abstract>
      <kwd-group>
        <kwd>nanoHUB</kwd>
        <kwd>HUBzero</kwd>
        <kwd>science gateways</kwd>
        <kwd>user behavior</kwd>
        <kwd>analytics</kwd>
        <kwd>cluster</kwd>
        <kwd>meander</kwd>
        <kwd>education</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>INTRODUCTION AND BACKGROUND
nanoHUB is a scientific knowledge platform that has
enabled over 3,500 researchers and educators to share 500+
research simulation tools and models as well as 6,000+ lectures
and tutorials globally through a novel cyberinfrastructure.
nanoHUB annually serves 17,000+ registered users with over 1
million simulations in an end-to-end user-oriented scientific
computing cloud. Over 1.5 million visitors access the openly
available web content items annually. These might be
considered impressive summative numbers, but they do not
address if the site has any impact or what these users are doing.</p>
      <p>Understanding these numbers requires some background on
the original intentions and cyberinfrastructure developments
around nanoHUB. Fundamental issues raised by
peerreviewers were the perceived ability of a University project to
Copyright © 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
provide a stable, national-level infrastructure, provide support
for the offered services, and provide compute cycles for an
ever-growing user base.</p>
      <p>
        From the very beginning in 1996 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the predecessor to
nanoHUB called Purdue Network Computing Hub (PUNCH)
was created to enable researchers to share their code without
re-writes through novel web interfaces with end-users in
education and research. PUNCH was so novel that even the
web-server had to be created within the team. By 2004 the
standard web-form-interfaces were antiquated and did not
inspire the interactive exploration of simulation results with
rapid “What If?” questions that users might have. Users had to
download their simulation data to manipulate them in a form
where they can be truly used. nanoHUB was not an
end-toend usage platform. It became clear that the system had to be
revamped to enable the hosting of user-friendly
engineeringuse inspired interactive applications. Such interactive sessions
had to be hosted in a reliable, scalable middleware that was
running in production mode, not as a research paper
demonstration. 3D dataset exploration had to be supported on
remote, dedicated GPUs that deliver the results to end users.
      </p>
      <p>RAPPTURE, the Rapid APPlication infrastrucTURE
toolkit [2] enabled researchers, who typically did not have any
graphical user interfaces to their codes to describe the input and
outputs of their codes in XML and to generate a GUI. New
middleware [3] enabled 1,000+ users to be hosted
simultaneously on a moderate cluster of about 20 compute
nodes. A novel remote GPU-based visualization system [4]
supported hundreds of simultaneous sessions. nanoHUB
established the first community accounts on TeraGrid and OSG
which would execute heavy-lifting nanoHUB simulation jobs
completely transparently on behalf of users who had no
accounts on these grid platforms [5]. We developed processes
[6] to continually test the reliability of these remote grid
services to ensure smooth user services. For application
support we developed policies and operational infrastructure
that enabled tool contributors to support and improve their
tools through question &amp; answer forums and through wishlists.
As this novel infrastructure emerged in 2005 we observed rapid
growth in the simulation user base from the historical numbers
of 500 annual users to over 10,000 in a few years. As
questions of technical feasibility were addressed new questions
as to actual and potential impact emerged.</p>
      <p>Early-on our peer reviewers raised fundamental questions
whether such research-based simulation tools could be used by
other researchers at all and if these tools could be used in
education without specific customizations. The nanoHUB
team developed analytics that documented nanoHUB use
research through reference and citation searches in the
scientific literature. Today we can document over 2,200 papers
that cite nanoHUB and we keep track of the used resources and
tools, to provide attribution to the published tools. When we
showed the first 200 formal citations our peers remained
unconvinced that this could be good research. We then began
to track secondary citations, which today sum to over 30,000
resulting in an h-index of 82.</p>
      <p>
        Our peers had a similarly strong opinion that research tools
could not be used in education. We therefore developed novel
clustering algorithms [7] that documented systematic used of
simulation tools in formal education settings. Today we can
show that over 35,000 students in over 1,800 classes at over
180 institutions have used nanoHUB in formalized education
settings. We could also measure the time-to-adoption between
tool publication and first-time systematic use in a classroom.
The median time was determined to be less than 6 months.
From the analysis of research use and education use we can
begin to qualify the attributes of the underlying simulation
tools. We found significant use in education and in research
for many of the nanoHUB tools. These research and education
impact studies are documented in detail in Nature
Nanotechnology [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ].
      </p>
      <p>We used retrospective data analytics to show how
simulation tools were used in structured education and how
simulation tools were used in novel research. We showed that
the transition from research tool publication to adoption in the
classroom is happening rapidly in typically less than six
months and demonstrated through longitudinal data how
research tools migrate into education. With the use of these
retrospective analytics, we have made strategic decisions in
terms of tool and content developments and justified continued
investments by NSF into nanoHUB.</p>
      <p>As we migrate towards a sustainable nanoHUB we must
embrace similar processes pursued by in similar platforms such
as Uber or AirBnB: we need to create actionable data
analytics that can rapidly support user experience and help
grow the supply in the two-sided market platform – we need to
improve the experience of providers as well as end-users.</p>
    </sec>
    <sec id="sec-2">
      <title>II RESEARCH QUESTIONS</title>
      <p>Beyond raw numbers of users and simulations, we have
over the years continued to ask ourselves: How do users
behave in their virtual world of a simulation tool? More
specifically:
- How do they “travel” through the design/exploration world?
- How many individual simulations do they run within one
session?
- How many parameters do users change?
- How different do researchers, classroom users, and
selfstudy users behave?
- How different do different classes behave?
- Does different class instruction material / scaffolding make a
difference?
- Can we provide feedback to instructors on their classrooms?
- Given certain usage patterns inside the tool:
- Can we improve the tools and provide feedback to the
developers?</p>
      <p>There are a variety of different requirements that need to be
met to address some of these questions in a scalable
infrastructure such as:
- Storage/availability of individual simulation runs within user
sessions
- A data description language that is shared across different
tools
- A large set of simulation runs and participants
- Other user data such as classroom participation, or researcher
identification, geolocation, etc.</p>
      <p>In the next Sections we describe some of our first results that
begin to address some of these questions.</p>
      <p>
        For our initial study presented here we focus on the user
behavior for PN Junction Lab [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ] which is consistently one of
the top 10 nanoHUB tools [10] within any year. Despite our
codename pntoy the tool is powered by an industrial strength
semiconductor device modeling tool called PADRE [
        <xref ref-type="bibr" rid="ref9">11</xref>
        ].
Instead of learning the complex PADRE input language that
involves gridding, geometry, material and environmental
specifications, users can easily ask “What if?” questions in a
toy-like fashion.
      </p>
    </sec>
    <sec id="sec-3">
      <title>II SEARCHERS AND WILDCATTERS</title>
      <p>RAPPTURE provides a rather generic description of
simulation tool inputs and outputs. Over 90% of the 500+
nanoHUB simulation tools utilize RAPPTURE as their data
description language. With existing simulation logs we can
now begin to study the user behavior inside simulation tools.
Each simulation tool typically consists of 10 to 50 parameters
that are exposed to the users. Most of these parameters are
freeform numbers such as length, doping, effective mass,
dielectric constant, temperature etc. with their specific units,
while there is also a significant set of discrete options such as
model or geometry choices. Assuming that each parameter
might have just 10 reasonable choices, then each tool spans a
configurational design space of at least 1010 to 1050. The
dimensionality of these tools is clearly too large to be
intuitively understood.</p>
      <p>
        We developed a visualization methodology [
        <xref ref-type="bibr" rid="ref10">12</xref>
        ] to flatten
an N-dimensional space into 2 dimensions. Figure 1 shows the
conceptual mapping and shows two significantly different user
behaviors. A searcher, who moves through the design space in
subsequent steps that appear to indicate a method or a goal. A
wildcatter who modifies, apparently wildly, the same set of
parameters and appears to jump throughout the design space.
Within the same publication we also documented the
development of a “Searchiness” index that assigns a single
value to the degree a user behaves like a prototypical wildcatter
(Searchiness=0) and prototypical searcher (Searchiness=1).
(c)
Figure 1: a) Visual representation of a multidimensional
space in two dimensions. b) a prototypical searcher. c) a
prototypical wildcatter.
      </p>
      <p>In this paper we show the analysis of a whole user
population using a specific tool and fuse that data set with
specific classroom users.</p>
    </sec>
    <sec id="sec-4">
      <title>III CLASSROOM CLUSTERS</title>
      <p>To demonstrate our ability to fuse different data sets from
our datastore we pick two different class clusters with
significantly different characteristics as depicted in Figure 2.
Class C12 is a class that reoccurred in 15 times between 2008
and 2018. We have pntoy simulation data from 7 classes within
2014 to 2018 for 109 users who ran 180 sessions. Historically,
we do not have the simulation data from all users in that time
frame. Going forward in the future we have developed a
simulation caching system where all Rappture simulations are
stored and users will receive stored solutions if they exist. The
cluster view in Figure 2 shows a subset, the individual class
held in the fall 2015 with 40 students who ran 80 sessions.
C12 only uses pntoy. Class C16 uses 6 different tools
throughout a semester. pntoy is one of these 6 tools used by 20
users in 29 sessions. The visual cluster representation in Figure
2 clearly shows the temporal behavior of 7 users who have
used all 6 tools in the class. In the next section we will
compare the behavior in these classes against all available data
and against all self-study users within the same region (Texas).</p>
    </sec>
    <sec id="sec-5">
      <title>IV USER BEHAVIOR DISTRIBUTIONS</title>
      <p>Figure 3 shows the Searchiness distribution of all 2,747
geolocated simulation sessions of pntoy by 1,865 users in the time
frame of 2014-2018. The complete distribution of all runs
shows clear peaks around 0 (wildcatters) and 1 (searchers).
Class cluster C12 is a subset of all the available data consistent
of 40 users with 80 sessions. This cluster usage uses only a
single tool in the whole class. Wildcatter behavior appears to
dominate this class C12. In contrast the smaller class that uses
in total 6 tools, including pntoy with 20 users and 28 sessions
shows a distribution that seems to indicate more searchers.</p>
      <p>Finally, we look at a third population within the users.
These are all the geo-located users in Texas (the location of
C12 and C16) who have not been identified as participants in
in any classes. We title this group of 20 who ran 36 simulation
sessions in pntoy as “self-study” users. These users show yet a
different distribution of Searchiness compared to the other
populations.</p>
      <p>Next to Searchiness, which is a computed model metric, we
can also look at a simple raw number, which is the number of
queries each individual has performed within a single tool
session. Within each tool session a user can execute the tool
multiple times and compare results as visualized in Figure 1.
Figure 4 shows the normalized distribution of queries executed
by the 4 different populations we examined in Figure 3. The
number of queries does not reveal much information except
that the overall population runs more queries than the 2 Texas
classes and the Texas self-study users. The classes and
selfstudy users show a rather strong drop-off for more than the
minimal queries of 4, which is the minimal number of queries
needed to define Searchiness. Initial analysis does not seem to
indicate a strong correlation to between Searchiness and
number of queries.</p>
    </sec>
    <sec id="sec-6">
      <title>VI CONCLUSION</title>
      <p>We report the development of a nanoHUB infrastructure
that begins to enable the study of user behavior in virtual
worlds of simulation tools. We use the previously published
model index Searchiness and compute it for a complete data set
of simulation sessions within a specific tool. We fuse data sets
of class cluster identification with the model index Searchiness
and number of queries. No surprising results are seen or
critical insight gained at this stage. We observe in the data
that different user populations appear to behave differently in
terms of Searchiness and classes seem to appear similar in
terms of number of queries. At this stage the data opens new
vectors for questions such as:
- Do all single-tool classes have similar behavior?
- Do classes with more diverse tool use or better scaffolding
foster more search-like behavior?
- Can similar behavior differences be seen with the other tools
that are used in classes?
- What does a peak in Searchiness value of 0.5 mean? Do we
need to refine the Searchiness index?
- Do we need to identify other behavioral metrics in addition
to Searchiness?
- Do the users who use other nanoHUB material outside the
tools behave differently than the ones that use tools only?
We conclude that this work is a first demonstrator that
indicates that we can assess the simulation behavior of different
user populations inside nanoHUB. We plan to refine these
metrics and classifiers to gain more insights on the user
behavior, and ultimately influence their behavior during use.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGEMENTS</title>
      <p>Funding by the US National Science Foundation under Grant
Nos. EEC-0228390, EEC-0634750, OCI-0438246,
OCI0721680, and EEC-1227110 as well as Purdue University is
gratefully acknowledged.
[10] https://nanohub.org/usage/tools provides nanoHUB tool listings ranked
by various criteria, such as number of users, number of simulations, wall
clock time, etc.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.H.</given-names>
            <surname>Kapadia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.A.B.</given-names>
            <surname>Fortes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            <surname>Lundstrom</surname>
          </string-name>
          , The Semiconductor Simulation Hub:
          <article-title>A network-based microelectronics simulation laboratory</article-title>
          ,
          <source>Proceedings of the Twelfth Biennial Conference:</source>
          [5]
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Klimeck</surname>
          </string-name>
          ,
          <string-name>
            <surname>Michael</surname>
            <given-names>McLennan</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Sean</given-names>
            <surname>Brophy</surname>
          </string-name>
          , George Adams III.,
          <string-name>
            <surname>Mark</surname>
            <given-names>Lundstrom</given-names>
          </string-name>
          ,
          <article-title>"nanoHUB.org: Advancing Education and Research in Nanotechnology"</article-title>
          ,
          <source>IEEE Computers in Engineering and Science (CISE)</source>
          , Vol.
          <volume>10</volume>
          ,
          <issue>Issue</issue>
          : 5,
          <string-name>
            <surname>Page</surname>
          </string-name>
          (s):
          <fpage>17</fpage>
          -
          <lpage>23</lpage>
          ,
          <string-name>
            <surname>Sept</surname>
          </string-name>
          .-Oct.
          <year>2008</year>
          ;doi:10.1109/
          <string-name>
            <surname>MCSE</surname>
          </string-name>
          .
          <year>2008</year>
          .120
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Lynn</given-names>
            <surname>Zentner</surname>
          </string-name>
          , Steven Clark, Krishna Madhavan, Swaroop Shivarajapura, Victoria Farnsworth, Gerhard Klimeck,
          <article-title>"Automated Grid-Probe System to Improve End-To-End Grid Reliability for a Science Gateway"</article-title>
          ,
          <source>Proceedings of TeraGrid 2011 conference. July 18- 21</source>
          ,
          <year>2011</year>
          ,
          <string-name>
            <given-names>Salt</given-names>
            <surname>Lake</surname>
          </string-name>
          <string-name>
            <surname>City</surname>
          </string-name>
          ,
          <source>ACM proceedings, ISBN: 978-1-4503-0888- 5;doi:10.1145/2016741</source>
          .2016789
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          University/Government/Industry Microelectronics Symposium,
          <year>1997</year>
          , IEEE Xplore, DOI: 10.1109/UGIM.
          <year>1997</year>
          .616686
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Michael McLennan</surname>
          </string-name>
          (
          <year>2005</year>
          ),
          <article-title>"Add Rappture to Your Software Development - Learning Module,"</article-title>
          https://nanohub.org/resources/240.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Mclennan</surname>
          </string-name>
          ,
          <string-name>
            <surname>Rick Kennell</surname>
          </string-name>
          ,
          <article-title>HUBzero: A Platform for Dissemination and Collaboration in Computational Science and Engineering</article-title>
          ,
          <source>IEEE Computing in Science and Engineering</source>
          <volume>12</volume>
          (
          <issue>2</issue>
          ):
          <fpage>48</fpage>
          -
          <lpage>53</lpage>
          ,
          <year>2010</year>
          DOI: 10.1109/
          <string-name>
            <surname>MCSE</surname>
          </string-name>
          .
          <year>2010</year>
          .41
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Wei</given-names>
            <surname>Qiao</surname>
          </string-name>
          ,
          <string-name>
            <surname>Michael</surname>
            <given-names>McLennan</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Rick</given-names>
            <surname>Kennel</surname>
          </string-name>
          , David Ebert,
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Klimeck</surname>
          </string-name>
          ,
          <article-title>"Hub-based Simulation and Graphics Hardware Accelerated Visualization for Nanotechnology Applications"</article-title>
          .
          <source>IEEE Transactions on Visualization and Computer Graphics</source>
          , Vol.
          <volume>12</volume>
          ,
          <issue>Issue</issue>
          : 5,
          <string-name>
            <surname>Page</surname>
          </string-name>
          (s):
          <fpage>1061</fpage>
          -
          <lpage>1068</lpage>
          ,
          <string-name>
            <surname>Sept</surname>
          </string-name>
          .-Oct.
          <year>2006</year>
          ;doi : 10.1109/TVCG.
          <year>2006</year>
          .150
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Krishna</given-names>
            <surname>Madhavan</surname>
          </string-name>
          , Michael Zentner, Gerhard Klimeck,
          <article-title>"Learning and research in the cloud"</article-title>
          ,
          <source>Nature Nanotechnology 8</source>
          ,
          <fpage>786</fpage>
          -
          <lpage>789</lpage>
          (
          <year>2013</year>
          ); doi:10.1038/nnano.
          <year>2013</year>
          .231
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Dragica</given-names>
            <surname>Vasileska</surname>
          </string-name>
          , Matteo Mannino,
          <string-name>
            <surname>Michael</surname>
            <given-names>McLennan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xufeng</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Gerhard Klimeck, Saumitra Raj Mehrotra, Benjamin P Haley (
          <year>2014</year>
          ),
          <article-title>"PN Junction Lab,"</article-title>
          https://nanohub.org/resources/pntoy. (DOI: 10.21981/D3GH9B95N).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Mark</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>Kent</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>Muhammad</given-names>
          </string-name>
          <string-name>
            <surname>Alam</surname>
          </string-name>
          , Steven Clark, Xufeng Wang, Gerhard Klimeck, Dragica
          <string-name>
            <surname>Vasileska</surname>
          </string-name>
          (
          <year>2014</year>
          ),
          <article-title>"</article-title>
          <string-name>
            <surname>Padre</surname>
          </string-name>
          ," https://nanohub.org/resources/padre. (DOI: 10.21981/D30C4SK7Z).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Nathan</surname>
            <given-names>Denny</given-names>
          </string-name>
          , Gerhard Klimeck,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Zentner</surname>
          </string-name>
          ,
          <article-title>"Visualizing User Interactions with Simulation Tool"</article-title>
          ,
          <source>Proceedings of the 13th Gateway Computing Environments Conference, September 25-27</source>
          ,
          <year>2018</year>
          , Austin, TX
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>