One-year test-retest reliability of ten vision tests in Canadian athletes

preprint OA: closed
Full text JSON View at publisher

Abstract

Background: : Vision tests are used in concussion management and baseline testing. Concussions, however, often occur months after baseline testing and reliability studies generally examine intervals limited to days or one week. Our objective was to determine the one-year test-retest reliability of these tests. Methods: : We assessed one-year test-retest reliability of ten vision tests in elite Canadian athletes followed by the Institut National du Sport du Quebec. We included athletes who completed two baseline (preseason) annual evaluations by one clinician within 365±30 days. We excluded athletes with any concussion or vision training in between the annual evaluations or presented with any factor that is believed to affect the tests (e.g. migraines). Data were collected from clinical charts. We evaluated test-retest reliability using Intraclass Correlation Coefficient (ICC) and 95% limits of agreement (LoA). Results: : We examined nine female and seven male athletes with a mean age of 22.7 (SD 4.5) years. Among the vision tests, we observed excellent test-retest reliability in Positive Fusional Vergence at 30cm (ICC=0.93) but this dropped to 0.53 when an outlier was excluded in a sensitivity analysis. There was good to moderate reliability in Negative Fusional Vergence at 30cm (ICC=0.78), Phoria at 30cm (ICC=0.68), Near Point of Convergence break (ICC=0.65) and Saccades (ICC=0.61). The ICC for Positive Fusional Vergence at 3m (ICC=0.56) also decreased to 0.45 after removing two outliers. We found poor reliability in Near Point of Convergence (ICC=0.47), Gross Stereoscopic Acuity (ICC=0.03) and Negative Fusional Vergence at 3m (ICC=0.0). ICC for Phoria at 3m was not appropriate because scores were identical in 14/16 athletes. 95% LoA of the majority of tests were ±40% to ±90%. Conclusions: : Five tests had good to moderate one-year test-retest reliability. The remaining tests had poor reliability. The tests would therefore be useful only if concussion has a moderate-large effect on scores.
Full text 304,809 characters · extracted from preprint-html · click to expand
One-year test-retest reliability of ten vision... | F1000Research "use strict";function _typeof(t){return(_typeof="function"==typeof Symbol&&"symbol"==typeof Symbol.iterator?function(t){return typeof t}:function(t){return t&&"function"==typeof Symbol&&t.constructor===Symbol&&t!==Symbol.prototype?"symbol":typeof t})(t)}!function(){var t=function(){var t,e,o=[],n=window,r=n;for(;r;){try{if(r.frames.__tcfapiLocator){t=r;break}}catch(t){}if(r===n.top)break;r=r.parent}t||(!function t(){var e=n.document,o=!!n.frames.__tcfapiLocator;if(!o)if(e.body){var r=e.createElement("iframe");r.style.cssText="display:none",r.name="__tcfapiLocator",e.body.appendChild(r)}else setTimeout(t,5);return!o}(),n.__tcfapi=function(){for(var t=arguments.length,n=new Array(t),r=0;r 3&&2===parseInt(n[1],10)&&"boolean"==typeof n[3]&&(e=n[3],"function"==typeof n[2]&&n[2]("set",!0)):"ping"===n[0]?"function"==typeof n[2]&&n[2]({gdprApplies:e,cmpLoaded:!1,cmpStatus:"stub"}):o.push(n)},n.addEventListener("message",(function(t){var e="string"==typeof t.data,o={};if(e)try{o=JSON.parse(t.data)}catch(t){}else o=t.data;var n="object"===_typeof(o)&&null!==o?o.__tcfapiCall:null;n&&window.__tcfapi(n.command,n.version,(function(o,r){var a={__tcfapiReturn:{returnValue:o,success:r,callId:n.callId}};t&&t.source&&t.source.postMessage&&t.source.postMessage(e?JSON.stringify(a):a,"*")}),n.parameter)}),!1))};"undefined"!=typeof module?module.exports=t:t()}(); dataLayer = dataLayer || []; // Standard GTM initialization - Google Consent Mode handles consent automatically (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start': new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0], j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src= 'https://www.googletagmanager.com/gtm.js?id='+i+dl+ '>m_auth=hzk0Vc3qFsQYhCrIoHz68A>m_preview=env-1>m_cookies_win=x';f.parentNode.insertBefore(j,f); })(window,document,'script','dataLayer','GTM-MWFK8L5J'); ;window.NREUM||(NREUM={});NREUM.init={distributed_tracing:{enabled:true},privacy:{cookies_enabled:true},ajax:{deny_list:["bam.nr-data.net"]}}; ;NREUM.loader_config={accountID:"438030",trustKey:"438030",agentID:"772317073",licenseKey:"97f8f67f26",applicationID:"772317073"} ;NREUM.info={beacon:"bam.nr-data.net",errorBeacon:"bam.nr-data.net",licenseKey:"97f8f67f26",applicationID:"772317073",sa:1} ;/*! For license information please see nr-loader-spa-1.236.0.min.js.LICENSE.txt */ (()=>{"use strict";var e,t,r={5763:(e,t,r)=>{r.d(t,{P_:()=>l,Mt:()=>g,C5:()=>s,DL:()=>v,OP:()=>T,lF:()=>D,Yu:()=>y,Dg:()=>h,CX:()=>c,GE:()=>b,sU:()=>_});var n=r(8632),i=r(9567);const o={beacon:n.ce.beacon,errorBeacon:n.ce.errorBeacon,licenseKey:void 0,applicationID:void 0,sa:void 0,queueTime:void 0,applicationTime:void 0,ttGuid:void 0,user:void 0,account:void 0,product:void 0,extra:void 0,jsAttributes:{},userAttributes:void 0,atts:void 0,transactionName:void 0,tNamePlain:void 0},a={};function s(e){if(!e)throw new Error("All info objects require an agent identifier!");if(!a[e])throw new Error("Info for ".concat(e," was never set"));return a[e]}function c(e,t){if(!e)throw new Error("All info objects require an agent identifier!");a[e]=(0,i.D)(t,o),(0,n.Qy)(e,a[e],"info")}var u=r(7056);const d=()=>{const e={blockSelector:"[data-nr-block]",maskInputOptions:{password:!0}};return{allow_bfcache:!0,privacy:{cookies_enabled:!0},ajax:{deny_list:void 0,enabled:!0,harvestTimeSeconds:10},distributed_tracing:{enabled:void 0,exclude_newrelic_header:void 0,cors_use_newrelic_header:void 0,cors_use_tracecontext_headers:void 0,allowed_origins:void 0},session:{domain:void 0,expiresMs:u.oD,inactiveMs:u.Hb},ssl:void 0,obfuscate:void 0,jserrors:{enabled:!0,harvestTimeSeconds:10},metrics:{enabled:!0},page_action:{enabled:!0,harvestTimeSeconds:30},page_view_event:{enabled:!0},page_view_timing:{enabled:!0,harvestTimeSeconds:30,long_task:!1},session_trace:{enabled:!0,harvestTimeSeconds:10},harvest:{tooManyRequestsDelay:60},session_replay:{enabled:!1,harvestTimeSeconds:60,sampleRate:.1,errorSampleRate:.1,maskTextSelector:"*",maskAllInputs:!0,get blockClass(){return"nr-block"},get ignoreClass(){return"nr-ignore"},get maskTextClass(){return"nr-mask"},get blockSelector(){return e.blockSelector},set blockSelector(t){e.blockSelector+=",".concat(t)},get maskInputOptions(){return e.maskInputOptions},set maskInputOptions(t){e.maskInputOptions={...t,password:!0}}},spa:{enabled:!0,harvestTimeSeconds:10}}},f={};function l(e){if(!e)throw new Error("All configuration objects require an agent identifier!");if(!f[e])throw new Error("Configuration for ".concat(e," was never set"));return f[e]}function h(e,t){if(!e)throw new Error("All configuration objects require an agent identifier!");f[e]=(0,i.D)(t,d()),(0,n.Qy)(e,f[e],"config")}function g(e,t){if(!e)throw new Error("All configuration objects require an agent identifier!");var r=l(e);if(r){for(var n=t.split("."),i=0;i {r.d(t,{D:()=>i});var n=r(50);function i(e,t){try{if(!e||"object"!=typeof e)return(0,n.Z)("Setting a Configurable requires an object as input");if(!t||"object"!=typeof t)return(0,n.Z)("Setting a Configurable requires a model to set its initial properties");const r=Object.create(Object.getPrototypeOf(t),Object.getOwnPropertyDescriptors(t)),o=0===Object.keys(r).length?e:r;for(let a in o)if(void 0!==e[a])try{"object"==typeof e[a]&&"object"==typeof t[a]?r[a]=i(e[a],t[a]):r[a]=e[a]}catch(e){(0,n.Z)("An error occurred while setting a property of a Configurable",e)}return r}catch(e){(0,n.Z)("An error occured while setting a Configurable",e)}}},6818:(e,t,r)=>{r.d(t,{Re:()=>i,gF:()=>o,q4:()=>n});const n="1.236.0",i="PROD",o="CDN"},385:(e,t,r)=>{r.d(t,{FN:()=>a,IF:()=>u,Nk:()=>f,Tt:()=>s,_A:()=>o,il:()=>n,pL:()=>c,v6:()=>i,w1:()=>d});const n="undefined"!=typeof window&&!!window.document,i="undefined"!=typeof WorkerGlobalScope&&("undefined"!=typeof self&&self instanceof WorkerGlobalScope&&self.navigator instanceof WorkerNavigator||"undefined"!=typeof globalThis&&globalThis instanceof WorkerGlobalScope&&globalThis.navigator instanceof WorkerNavigator),o=n?window:"undefined"!=typeof WorkerGlobalScope&&("undefined"!=typeof self&&self instanceof WorkerGlobalScope&&self||"undefined"!=typeof globalThis&&globalThis instanceof WorkerGlobalScope&&globalThis),a=""+o?.location,s=/iPad|iPhone|iPod/.test(navigator.userAgent),c=s&&"undefined"==typeof SharedWorker,u=(()=>{const e=navigator.userAgent.match(/Firefox[/\s](\d+\.\d+)/);return Array.isArray(e)&&e.length>=2?+e[1]:0})(),d=Boolean(n&&window.document.documentMode),f=!!navigator.sendBeacon},1117:(e,t,r)=>{r.d(t,{w:()=>o});var n=r(50);const i={agentIdentifier:"",ee:void 0};class o{constructor(e){try{if("object"!=typeof e)return(0,n.Z)("shared context requires an object as input");this.sharedContext={},Object.assign(this.sharedContext,i),Object.entries(e).forEach((e=>{let[t,r]=e;Object.keys(i).includes(t)&&(this.sharedContext[t]=r)}))}catch(e){(0,n.Z)("An error occured while setting SharedContext",e)}}}},8e3:(e,t,r)=>{r.d(t,{L:()=>d,R:()=>c});var n=r(2177),i=r(1284),o=r(4322),a=r(3325);const s={};function c(e,t){const r={staged:!1,priority:a.p[t]||0};u(e),s[e].get(t)||s[e].set(t,r)}function u(e){e&&(s[e]||(s[e]=new Map))}function d(){let e=arguments.length>0&&void 0!==arguments[0]?arguments[0]:"",t=arguments.length>1&&void 0!==arguments[1]?arguments[1]:"feature";if(u(e),!e||!s[e].get(t))return a(t);s[e].get(t).staged=!0;const r=[...s[e]];function a(t){const r=e?n.ee.get(e):n.ee,a=o.X.handlers;if(r.backlog&&a){var s=r.backlog[t],c=a[t];if(c){for(var u=0;s&&u {let[t,r]=e;return r.staged}))&&(r.sort(((e,t)=>e[1].priority-t[1].priority)),r.forEach((e=>{let[t]=e;a(t)})))}function f(e,t){var r=e[1];(0,i.D)(t[r],(function(t,r){var n=e[0];if(r[0]===n){var i=r[1],o=e[3],a=e[2];i.apply(o,a)}}))}},2177:(e,t,r)=>{r.d(t,{c:()=>f,ee:()=>u});var n=r(8632),i=r(2210),o=r(1284),a=r(5763),s="nr@context";let c=(0,n.fP)();var u;function d(){}function f(e){return(0,i.X)(e,s,l)}function l(){return new d}function h(){u.aborted=!0,u.backlog={}}c.ee?u=c.ee:(u=function e(t,r){var n={},c={},f={},g=!1;try{g=16===r.length&&(0,a.OP)(r).isolatedBacklog}catch(e){}var p={on:b,addEventListener:b,removeEventListener:y,emit:v,get:x,listeners:w,context:m,buffer:A,abort:h,aborted:!1,isBuffering:E,debugId:r,backlog:g?{}:t&&"object"==typeof t.backlog?t.backlog:{}};return p;function m(e){return e&&e instanceof d?e:e?(0,i.X)(e,s,l):l()}function v(e,r,n,i,o){if(!1!==o&&(o=!0),!u.aborted||i){t&&o&&t.emit(e,r,n);for(var a=m(n),s=w(e),d=s.length,f=0;fn,p:()=>i});var n=r(2177).ee.get("handle");function i(e,t,r,i,o){o?(o.buffer([e],i),o.emit(e,t,r)):(n.buffer([e],i),n.emit(e,t,r))}},4322:(e,t,r)=>{r.d(t,{X:()=>o});var n=r(5546);o.on=a;var i=o.handlers={};function o(e,t,r,o){a(o||n.E,i,e,t,r)}function a(e,t,r,i,o){o||(o="feature"),e||(e=n.E);var a=t[o]=t[o]||{};(a[r]=a[r]||[]).push([e,i])}},3239:(e,t,r)=>{r.d(t,{bP:()=>s,iz:()=>c,m$:()=>a});var n=r(385);let i=!1,o=!1;try{const e={get passive(){return i=!0,!1},get signal(){return o=!0,!1}};n._A.addEventListener("test",null,e),n._A.removeEventListener("test",null,e)}catch(e){}function a(e,t){return i||o?{capture:!!e,passive:i,signal:t}:!!e}function s(e,t){let r=arguments.length>2&&void 0!==arguments[2]&&arguments[2],n=arguments.length>3?arguments[3]:void 0;window.addEventListener(e,t,a(r,n))}function c(e,t){let r=arguments.length>2&&void 0!==arguments[2]&&arguments[2],n=arguments.length>3?arguments[3]:void 0;document.addEventListener(e,t,a(r,n))}},4402:(e,t,r)=>{r.d(t,{Ht:()=>u,M:()=>c,Rl:()=>a,ky:()=>s});var n=r(385);const i="xxxxxxxx-xxxx-4xxx-yxxx-xxxxxxxxxxxx";function o(e,t){return e?15&e[t]:16*Math.random()|0}function a(){const e=n._A?.crypto||n._A?.msCrypto;let t,r=0;return e&&e.getRandomValues&&(t=e.getRandomValues(new Uint8Array(31))),i.split("").map((e=>"x"===e?o(t,++r).toString(16):"y"===e?(3&o()|8).toString(16):e)).join("")}function s(e){const t=n._A?.crypto||n._A?.msCrypto;let r,i=0;t&&t.getRandomValues&&(r=t.getRandomValues(new Uint8Array(31)));const a=[];for(var s=0;s {r.d(t,{Bq:()=>n,Hb:()=>o,oD:()=>i});const n="NRBA",i=144e5,o=18e5},7894:(e,t,r)=>{function n(){return Math.round(performance.now())}r.d(t,{z:()=>n})},7243:(e,t,r)=>{r.d(t,{e:()=>o});var n=r(385),i={};function o(e){if(e in i)return i[e];if(0===(e||"").indexOf("data:"))return{protocol:"data"};let t;var r=n._A?.location,o={};if(n.il)t=document.createElement("a"),t.href=e;else try{t=new URL(e,r.href)}catch(e){return o}o.port=t.port;var a=t.href.split("://");!o.port&&a[1]&&(o.port=a[1].split("/")[0].split("@").pop().split(":")[1]),o.port&&"0"!==o.port||(o.port="https"===a[0]?"443":"80"),o.hostname=t.hostname||r.hostname,o.pathname=t.pathname,o.protocol=a[0],"/"!==o.pathname.charAt(0)&&(o.pathname="/"+o.pathname);var s=!t.protocol||":"===t.protocol||t.protocol===r.protocol,c=t.hostname===r.hostname&&t.port===r.port;return o.sameOrigin=s&&(!t.hostname||c),"/"===o.pathname&&(i[e]=o),o}},50:(e,t,r)=>{function n(e,t){"function"==typeof console.warn&&(console.warn("New Relic: ".concat(e)),t&&console.warn(t))}r.d(t,{Z:()=>n})},2587:(e,t,r)=>{r.d(t,{N:()=>c,T:()=>u});var n=r(2177),i=r(5546),o=r(8e3),a=r(3325);const s={stn:[a.D.sessionTrace],err:[a.D.jserrors,a.D.metrics],ins:[a.D.pageAction],spa:[a.D.spa],sr:[a.D.sessionReplay,a.D.sessionTrace]};function c(e,t){const r=n.ee.get(t);e&&"object"==typeof e&&(Object.entries(e).forEach((e=>{let[t,n]=e;void 0===u[t]&&(s[t]?s[t].forEach((e=>{n?(0,i.p)("feat-"+t,[],void 0,e,r):(0,i.p)("block-"+t,[],void 0,e,r),(0,i.p)("rumresp-"+t,[Boolean(n)],void 0,e,r)})):n&&(0,i.p)("feat-"+t,[],void 0,void 0,r),u[t]=Boolean(n))})),Object.keys(s).forEach((e=>{void 0===u[e]&&(s[e]?.forEach((t=>(0,i.p)("rumresp-"+e,[!1],void 0,t,r))),u[e]=!1)})),(0,o.L)(t,a.D.pageViewEvent))}const u={}},2210:(e,t,r)=>{r.d(t,{X:()=>i});var n=Object.prototype.hasOwnProperty;function i(e,t,r){if(n.call(e,t))return e[t];var i=r();if(Object.defineProperty&&Object.keys)try{return Object.defineProperty(e,t,{value:i,writable:!0,enumerable:!1}),i}catch(e){}return e[t]=i,i}},1284:(e,t,r)=>{r.d(t,{D:()=>n});const n=(e,t)=>Object.entries(e||{}).map((e=>{let[r,n]=e;return t(r,n)}))},4351:(e,t,r)=>{r.d(t,{P:()=>o});var n=r(2177);const i=()=>{const e=new WeakSet;return(t,r)=>{if("object"==typeof r&&null!==r){if(e.has(r))return;e.add(r)}return r}};function o(e){try{return JSON.stringify(e,i())}catch(e){try{n.ee.emit("internal-error",[e])}catch(e){}}}},3960:(e,t,r)=>{r.d(t,{K:()=>a,b:()=>o});var n=r(3239);function i(){return"undefined"==typeof document||"complete"===document.readyState}function o(e,t){if(i())return e();(0,n.bP)("load",e,t)}function a(e){if(i())return e();(0,n.iz)("DOMContentLoaded",e)}},8632:(e,t,r)=>{r.d(t,{EZ:()=>u,Qy:()=>c,ce:()=>o,fP:()=>a,gG:()=>d,mF:()=>s});var n=r(7894),i=r(385);const o={beacon:"bam.nr-data.net",errorBeacon:"bam.nr-data.net"};function a(){return i._A.NREUM||(i._A.NREUM={}),void 0===i._A.newrelic&&(i._A.newrelic=i._A.NREUM),i._A.NREUM}function s(){let e=a();return e.o||(e.o={ST:i._A.setTimeout,SI:i._A.setImmediate,CT:i._A.clearTimeout,XHR:i._A.XMLHttpRequest,REQ:i._A.Request,EV:i._A.Event,PR:i._A.Promise,MO:i._A.MutationObserver,FETCH:i._A.fetch}),e}function c(e,t,r){let i=a();const o=i.initializedAgents||{},s=o[e]||{};return Object.keys(s).length||(s.initializedAt={ms:(0,n.z)(),date:new Date}),i.initializedAgents={...o,[e]:{...s,[r]:t}},i}function u(e,t){a()[e]=t}function d(){return function(){let e=a();const t=e.info||{};e.info={beacon:o.beacon,errorBeacon:o.errorBeacon,...t}}(),function(){let e=a();const t=e.init||{};e.init={...t}}(),s(),function(){let e=a();const t=e.loader_config||{};e.loader_config={...t}}(),a()}},7956:(e,t,r)=>{r.d(t,{N:()=>i});var n=r(3239);function i(e){let t=arguments.length>1&&void 0!==arguments[1]&&arguments[1],r=arguments.length>2?arguments[2]:void 0,i=arguments.length>3?arguments[3]:void 0;return void(0,n.iz)("visibilitychange",(function(){if(t)return void("hidden"==document.visibilityState&&e());e(document.visibilityState)}),r,i)}},1214:(e,t,r)=>{r.d(t,{em:()=>v,u5:()=>N,QU:()=>S,_L:()=>I,Gm:()=>L,Lg:()=>M,gy:()=>U,BV:()=>Q,Kf:()=>ee});var n=r(2177);const i="nr@original";var o=Object.prototype.hasOwnProperty,a=!1;function s(e,t){return e||(e=n.ee),r.inPlace=function(e,t,n,i,o){n||(n="");var a,s,c,u="-"===n.charAt(0);for(c=0;c 2?n-2:0),o=2;o {r(A[T],e,w),r(E[T],e,w)})),r(l._A,"fetch",y),t.on(y+"end",(function(e,r){var n=this;if(r){var i=r.headers.get("content-length");null!==i&&(n.rxSize=i),t.emit(y+"done",[null,r],n)}else t.emit(y+"done",[e],n)})),t}const O={},j=["pushState","replaceState"];function S(e){const t=function(e){return(e||n.ee).get("history")}(e);return!l.il||O[t.debugId]++||(O[t.debugId]=1,s(t).inPlace(window.history,j,"-")),t}var P=r(3239);const C={},R=["appendChild","insertBefore","replaceChild"];function I(e){const t=function(e){return(e||n.ee).get("jsonp")}(e);if(!l.il||C[t.debugId])return t;C[t.debugId]=!0;var r=s(t),i=/[?&](?:callback|cb)=([^&#]+)/,o=/(.*)\.([^.]+)/,a=/^(\w+)(\.|$)(.*)$/;function c(e,t){var r=e.match(a),n=r[1],i=r[3];return i?c(i,t[n]):t[n]}return r.inPlace(Node.prototype,R,"dom-"),t.on("dom-start",(function(e){!function(e){if(!e||"string"!=typeof e.nodeName||"script"!==e.nodeName.toLowerCase())return;if("function"!=typeof e.addEventListener)return;var n=(a=e.src,s=a.match(i),s?s[1]:null);var a,s;if(!n)return;var u=function(e){var t=e.match(o);if(t&&t.length>=3)return{key:t[2],parent:c(t[1],window)};return{key:e,parent:window}}(n);if("function"!=typeof u.parent[u.key])return;var d={};function f(){t.emit("jsonp-end",[],d),e.removeEventListener("load",f,(0,P.m$)(!1)),e.removeEventListener("error",l,(0,P.m$)(!1))}function l(){t.emit("jsonp-error",[],d),t.emit("jsonp-end",[],d),e.removeEventListener("load",f,(0,P.m$)(!1)),e.removeEventListener("error",l,(0,P.m$)(!1))}r.inPlace(u.parent,[u.key],"cb-",d),e.addEventListener("load",f,(0,P.m$)(!1)),e.addEventListener("error",l,(0,P.m$)(!1)),t.emit("new-jsonp",[e.src],d)}(e[0])})),t}var k=r(5763);const H={};function L(e){const t=function(e){return(e||n.ee).get("mutation")}(e);if(!l.il||H[t.debugId])return t;H[t.debugId]=!0;var r=s(t),i=k.Yu.MO;return i&&(window.MutationObserver=function(e){return this instanceof i?new i(r(e,"fn-")):i.apply(this,arguments)},MutationObserver.prototype=i.prototype),t}const z={};function M(e){const t=function(e){return(e||n.ee).get("promise")}(e);if(z[t.debugId])return t;z[t.debugId]=!0;var r=n.c,o=s(t),a=k.Yu.PR;return a&&function(){function e(r){var n=t.context(),i=o(r,"executor-",n,null,!1);const s=Reflect.construct(a,[i],e);return t.context(s).getCtx=function(){return n},s}l._A.Promise=e,Object.defineProperty(e,"name",{value:"Promise"}),e.toString=function(){return a.toString()},Object.setPrototypeOf(e,a),["all","race"].forEach((function(r){const n=a[r];e[r]=function(e){let i=!1;[...e||[]].forEach((e=>{this.resolve(e).then(a("all"===r),a(!1))}));const o=n.apply(this,arguments);return o;function a(e){return function(){t.emit("propagate",[null,!i],o,!1,!1),i=i||!e}}}})),["resolve","reject"].forEach((function(r){const n=a[r];e[r]=function(e){const r=n.apply(this,arguments);return e!==r&&t.emit("propagate",[e,!0],r,!1,!1),r}})),e.prototype=a.prototype;const n=a.prototype.then;a.prototype.then=function(){var e=this,i=r(e);i.promise=e;for(var a=arguments.length,s=new Array(a),c=0;c e())),t};function m(e,t){i.inPlace(t,["onreadystatechange"],"fn-",E)}function b(){var e=this,t=r.context(e);e.readyState>3&&!t.resolved&&(t.resolved=!0,r.emit("xhr-resolved",[],e)),i.inPlace(e,f,"fn-",E)}if(function(e,t){for(var r in e)t[r]=e[r]}(o,p),p.prototype=o.prototype,i.inPlace(p.prototype,J,"-xhr-",E),r.on("send-xhr-start",(function(e,t){m(e,t),function(e){h.push(e),a&&(y?y.then(A):u?u(A):(w=-w,x.data=w))}(t)})),r.on("open-xhr-start",m),a){var y=c&&c.resolve();if(!u&&!c){var w=1,x=document.createTextNode(w);new a(A).observe(x,{characterData:!0})}}else t.on("fn-end",(function(e){e[0]&&e[0].type===d||A()}));function A(){for(var e=0;e {r.d(t,{t:()=>n});const n=r(3325).D.ajax},6660:(e,t,r)=>{r.d(t,{A:()=>i,t:()=>n});const n=r(3325).D.jserrors,i="nr@seenError"},3081:(e,t,r)=>{r.d(t,{gF:()=>o,mY:()=>i,t9:()=>n,vz:()=>s,xS:()=>a});const n=r(3325).D.metrics,i="sm",o="cm",a="storeSupportabilityMetrics",s="storeEventMetrics"},4649:(e,t,r)=>{r.d(t,{t:()=>n});const n=r(3325).D.pageAction},7633:(e,t,r)=>{r.d(t,{Dz:()=>i,OJ:()=>a,qw:()=>o,t9:()=>n});const n=r(3325).D.pageViewEvent,i="firstbyte",o="domcontent",a="windowload"},9251:(e,t,r)=>{r.d(t,{t:()=>n});const n=r(3325).D.pageViewTiming},3614:(e,t,r)=>{r.d(t,{BST_RESOURCE:()=>i,END:()=>s,FEATURE_NAME:()=>n,FN_END:()=>u,FN_START:()=>c,PUSH_STATE:()=>d,RESOURCE:()=>o,START:()=>a});const n=r(3325).D.sessionTrace,i="bstResource",o="resource",a="-start",s="-end",c="fn"+a,u="fn"+s,d="pushState"},7836:(e,t,r)=>{r.d(t,{BODY:()=>A,CB_END:()=>E,CB_START:()=>u,END:()=>x,FEATURE_NAME:()=>i,FETCH:()=>_,FETCH_BODY:()=>v,FETCH_DONE:()=>m,FETCH_START:()=>p,FN_END:()=>c,FN_START:()=>s,INTERACTION:()=>l,INTERACTION_API:()=>d,INTERACTION_EVENTS:()=>o,JSONP_END:()=>b,JSONP_NODE:()=>g,JS_TIME:()=>T,MAX_TIMER_BUDGET:()=>a,REMAINING:()=>f,SPA_NODE:()=>h,START:()=>w,originalSetTimeout:()=>y});var n=r(5763);const i=r(3325).D.spa,o=["click","submit","keypress","keydown","keyup","change"],a=999,s="fn-start",c="fn-end",u="cb-start",d="api-ixn-",f="remaining",l="interaction",h="spaNode",g="jsonpNode",p="fetch-start",m="fetch-done",v="fetch-body-",b="jsonp-end",y=n.Yu.ST,w="-start",x="-end",A="-body",E="cb"+x,T="jsTime",_="fetch"},5938:(e,t,r)=>{r.d(t,{W:()=>o});var n=r(5763),i=r(2177);class o{constructor(e,t,r){this.agentIdentifier=e,this.aggregator=t,this.ee=i.ee.get(e,(0,n.OP)(this.agentIdentifier).isolatedBacklog),this.featureName=r,this.blocked=!1}}},9144:(e,t,r)=>{r.d(t,{j:()=>m});var n=r(3325),i=r(5763),o=r(5546),a=r(2177),s=r(7894),c=r(8e3),u=r(3960),d=r(385),f=r(50),l=r(3081),h=r(8632);function g(){const e=(0,h.gG)();["setErrorHandler","finished","addToTrace","inlineHit","addRelease","addPageAction","setCurrentRouteName","setPageViewName","setCustomAttribute","interaction","noticeError","setUserId"].forEach((t=>{e[t]=function(){for(var r=arguments.length,n=new Array(r),i=0;i 1?r-1:0),i=1;i {e.exposed&&e.api[t]&&o.push(e.api[t](...n))})),o.length>1?o:o[0]}(t,...n)}}))}var p=r(2587);function m(e){let t=arguments.length>1&&void 0!==arguments[1]?arguments[1]:{},m=arguments.length>2?arguments[2]:void 0,v=arguments.length>3?arguments[3]:void 0,{init:b,info:y,loader_config:w,runtime:x={loaderType:m},exposed:A=!0}=t;const E=(0,h.gG)();y||(b=E.init,y=E.info,w=E.loader_config),(0,i.Dg)(e,b||{}),(0,i.GE)(e,w||{}),(0,i.sU)(e,x),y.jsAttributes??={},d.v6&&(y.jsAttributes.isWorker=!0),(0,i.CX)(e,y),g();const T=function(e,t){t||(0,c.R)(e,"api");const h={};var g=a.ee.get(e),p=g.get("tracer"),m="api-",v=m+"ixn-";function b(t,r,n,o){const a=(0,i.C5)(e);return null===r?delete a.jsAttributes[t]:(0,i.CX)(e,{...a,jsAttributes:{...a.jsAttributes,[t]:r}}),x(m,n,!0,o||null===r?"session":void 0)(t,r)}function y(){}["setErrorHandler","finished","addToTrace","inlineHit","addRelease"].forEach((e=>h[e]=x(m,e,!0,"api"))),h.addPageAction=x(m,"addPageAction",!0,n.D.pageAction),h.setCurrentRouteName=x(m,"routeName",!0,n.D.spa),h.setPageViewName=function(t,r){if("string"==typeof t)return"/"!==t.charAt(0)&&(t="/"+t),(0,i.OP)(e).customTransaction=(r||"http://custom.transaction")+t,x(m,"setPageViewName",!0)()},h.setCustomAttribute=function(e,t){let r=arguments.length>2&&void 0!==arguments[2]&&arguments[2];if("string"==typeof e){if(["string","number"].includes(typeof t)||null===t)return b(e,t,"setCustomAttribute",r);(0,f.Z)("Failed to execute setCustomAttribute.\nNon-null value must be a string or number type, but a type of was provided."))}else(0,f.Z)("Failed to execute setCustomAttribute.\nName must be a string type, but a type of was provided."))},h.setUserId=function(e){if("string"==typeof e||null===e)return b("enduser.id",e,"setUserId",!0);(0,f.Z)("Failed to execute setUserId.\nNon-null value must be a string type, but a type of was provided."))},h.interaction=function(){return(new y).get()};var w=y.prototype={createTracer:function(e,t){var r={},i=this,a="function"==typeof t;return(0,o.p)(v+"tracer",[(0,s.z)(),e,r],i,n.D.spa,g),function(){if(p.emit((a?"":"no-")+"fn-start",[(0,s.z)(),i,a],r),a)try{return t.apply(this,arguments)}catch(e){throw p.emit("fn-err",[arguments,this,"string"==typeof e?new Error(e):e],r),e}finally{p.emit("fn-end",[(0,s.z)()],r)}}}};function x(e,t,r,i){return function(){return(0,o.p)(l.xS,["API/"+t+"/called"],void 0,n.D.metrics,g),i&&(0,o.p)(e+t,[(0,s.z)(),...arguments],r?null:this,i,g),r?void 0:this}}function A(){r.e(439).then(r.bind(r,7438)).then((t=>{let{setAPI:r}=t;r(e),(0,c.L)(e,"api")})).catch((()=>(0,f.Z)("Downloading runtime APIs failed...")))}return["actionText","setName","setAttribute","save","ignore","onEnd","getContext","end","get"].forEach((e=>{w[e]=x(v,e,void 0,n.D.spa)})),h.noticeError=function(e,t){"string"==typeof e&&(e=new Error(e)),(0,o.p)(l.xS,["API/noticeError/called"],void 0,n.D.metrics,g),(0,o.p)("err",[e,(0,s.z)(),!1,t],void 0,n.D.jserrors,g)},d.il?(0,u.b)((()=>A()),!0):A(),h}(e,v);return(0,h.Qy)(e,T,"api"),(0,h.Qy)(e,A,"exposed"),(0,h.EZ)("activatedFeatures",p.T),T}},3325:(e,t,r)=>{r.d(t,{D:()=>n,p:()=>i});const n={ajax:"ajax",jserrors:"jserrors",metrics:"metrics",pageAction:"page_action",pageViewEvent:"page_view_event",pageViewTiming:"page_view_timing",sessionReplay:"session_replay",sessionTrace:"session_trace",spa:"spa"},i={[n.pageViewEvent]:1,[n.pageViewTiming]:2,[n.metrics]:3,[n.jserrors]:4,[n.ajax]:5,[n.sessionTrace]:6,[n.pageAction]:7,[n.spa]:8,[n.sessionReplay]:9}}},n={};function i(e){var t=n[e];if(void 0!==t)return t.exports;var o=n[e]={exports:{}};return r[e](o,o.exports,i),o.exports}i.m=r,i.d=(e,t)=>{for(var r in t)i.o(t,r)&&!i.o(e,r)&&Object.defineProperty(e,r,{enumerable:!0,get:t[r]})},i.f={},i.e=e=>Promise.all(Object.keys(i.f).reduce(((t,r)=>(i.f[r](e,t),t)),[])),i.u=e=>(({78:"page_action-aggregate",147:"metrics-aggregate",242:"session-manager",317:"jserrors-aggregate",348:"page_view_timing-aggregate",412:"lazy-feature-loader",439:"async-api",538:"recorder",590:"session_replay-aggregate",675:"compressor",733:"session_trace-aggregate",786:"page_view_event-aggregate",873:"spa-aggregate",898:"ajax-aggregate"}[e]||e)+"."+{78:"ac76d497",147:"3dc53903",148:"1a20d5fe",242:"2a64278a",317:"49e41428",348:"bd6de33a",412:"2f55ce66",439:"30bd804e",538:"1b18459f",590:"cf0efb30",675:"ae9f91a8",733:"83105561",786:"06482edd",860:"03a8b7a5",873:"e6b09d52",898:"998ef92b"}[e]+"-1.236.0.min.js"),i.o=(e,t)=>Object.prototype.hasOwnProperty.call(e,t),e={},t="NRBA:",i.l=(r,n,o,a)=>{if(e[r])e[r].push(n);else{var s,c;if(void 0!==o)for(var u=document.getElementsByTagName("script"),d=0;d {s.onerror=s.onload=null,clearTimeout(h);var i=e[r];if(delete e[r],s.parentNode&&s.parentNode.removeChild(s),i&&i.forEach((e=>e(n))),t)return t(n)},h=setTimeout(l.bind(null,void 0,{type:"timeout",target:s}),12e4);s.onerror=l.bind(null,s.onerror),s.onload=l.bind(null,s.onload),c&&document.head.appendChild(s)}},i.r=e=>{"undefined"!=typeof Symbol&&Symbol.toStringTag&&Object.defineProperty(e,Symbol.toStringTag,{value:"Module"}),Object.defineProperty(e,"__esModule",{value:!0})},i.j=364,i.p="https://js-agent.newrelic.com/",(()=>{var e={364:0,953:0};i.f.j=(t,r)=>{var n=i.o(e,t)?e[t]:void 0;if(0!==n)if(n)r.push(n[2]);else{var o=new Promise(((r,i)=>n=e[t]=[r,i]));r.push(n[2]=o);var a=i.p+i.u(t),s=new Error;i.l(a,(r=>{if(i.o(e,t)&&(0!==(n=e[t])&&(e[t]=void 0),n)){var o=r&&("load"===r.type?"missing":r.type),a=r&&r.target&&r.target.src;s.message="Loading chunk "+t+" failed.\n("+o+": "+a+")",s.name="ChunkLoadError",s.type=o,s.request=a,n[1](s)}}),"chunk-"+t,t)}};var t=(t,r)=>{var n,o,[a,s,c]=r,u=0;if(a.some((t=>0!==e[t]))){for(n in s)i.o(s,n)&&(i.m[n]=s[n]);if(c)c(i)}for(t&&t(r);u {i.r(o);var e=i(3325),t=i(5763);const r=Object.values(e.D);function n(e){const n={};return r.forEach((r=>{n[r]=function(e,r){return!1!==(0,t.Mt)(r,"".concat(e,".enabled"))}(r,e)})),n}var a=i(9144);var s=i(5546),c=i(385),u=i(8e3),d=i(5938),f=i(3960),l=i(50);class h extends d.W{constructor(e,t,r){let n=!(arguments.length>3&&void 0!==arguments[3])||arguments[3];super(e,t,r),this.auto=n,this.abortHandler,this.featAggregate,this.onAggregateImported,n&&(0,u.R)(e,r)}importAggregator(){let e=arguments.length>0&&void 0!==arguments[0]?arguments[0]:{};if(this.featAggregate||!this.auto)return;const r=c.il&&!0===(0,t.Mt)(this.agentIdentifier,"privacy.cookies_enabled");let n;this.onAggregateImported=new Promise((e=>{n=e}));const o=async()=>{let t;try{if(r){const{setupAgentSession:e}=await Promise.all([i.e(860),i.e(242)]).then(i.bind(i,3228));t=e(this.agentIdentifier)}}catch(e){(0,l.Z)("A problem occurred when starting up session manager. This page will not start or extend any session.",e)}try{if(!this.shouldImportAgg(this.featureName,t))return void(0,u.L)(this.agentIdentifier,this.featureName);const{lazyFeatureLoader:r}=await i.e(412).then(i.bind(i,8582)),{Aggregate:o}=await r(this.featureName,"aggregate");this.featAggregate=new o(this.agentIdentifier,this.aggregator,e),n(!0)}catch(e){(0,l.Z)("Downloading and initializing ".concat(this.featureName," failed..."),e),this.abortHandler?.(),n(!1)}};c.il?(0,f.b)((()=>o()),!0):o()}shouldImportAgg(r,n){return r!==e.D.sessionReplay||!1!==(0,t.Mt)(this.agentIdentifier,"session_trace.enabled")&&(!!n?.isNew||!!n?.state.sessionReplay)}}var g=i(7633),p=i(7894);class m extends h{static featureName=g.t9;constructor(r,n){let i=!(arguments.length>2&&void 0!==arguments[2])||arguments[2];if(super(r,n,g.t9,i),("undefined"==typeof PerformanceNavigationTiming||c.Tt)&&"undefined"!=typeof PerformanceTiming){const n=(0,t.OP)(r);n[g.Dz]=Math.max(Date.now()-n.offset,0),(0,f.K)((()=>n[g.qw]=Math.max((0,p.z)()-n[g.Dz],0))),(0,f.b)((()=>{const t=(0,p.z)();n[g.OJ]=Math.max(t-n[g.Dz],0),(0,s.p)("timing",["load",t],void 0,e.D.pageViewTiming,this.ee)}))}this.importAggregator()}}var v=i(1117),b=i(1284);class y extends v.w{constructor(e){super(e),this.aggregatedData={}}store(e,t,r,n,i){var o=this.getBucket(e,t,r,i);return o.metrics=function(e,t){t||(t={count:0});return t.count+=1,(0,b.D)(e,(function(e,r){t[e]=w(r,t[e])})),t}(n,o.metrics),o}merge(e,t,r,n,i){var o=this.getBucket(e,t,n,i);if(o.metrics){var a=o.metrics;a.count+=r.count,(0,b.D)(r,(function(e,t){if("count"!==e){var n=a[e],i=r[e];i&&!i.c?a[e]=w(i.t,n):a[e]=function(e,t){if(!t)return e;t.c||(t=x(t.t));return t.min=Math.min(e.min,t.min),t.max=Math.max(e.max,t.max),t.t+=e.t,t.sos+=e.sos,t.c+=e.c,t}(i,a[e])}}))}else o.metrics=r}storeMetric(e,t,r,n){var i=this.getBucket(e,t,r);return i.stats=w(n,i.stats),i}getBucket(e,t,r,n){this.aggregatedData[e]||(this.aggregatedData[e]={});var i=this.aggregatedData[e][t];return i||(i=this.aggregatedData[e][t]={params:r||{}},n&&(i.custom=n)),i}get(e,t){return t?this.aggregatedData[e]&&this.aggregatedData[e][t]:this.aggregatedData[e]}take(e){for(var t={},r="",n=!1,i=0;i t.max&&(t.max=e),e 2&&void 0!==arguments[2])||arguments[2];super(e,r,j.t,n),c.il&&((0,t.OP)(e).initHidden=Boolean("hidden"===document.visibilityState),(0,N.N)((()=>(0,s.p)("docHidden",[(0,p.z)()],void 0,j.t,this.ee)),!0),(0,O.bP)("pagehide",(()=>(0,s.p)("winPagehide",[(0,p.z)()],void 0,j.t,this.ee))),this.importAggregator())}}var P=i(3081);class C extends h{static featureName=P.t9;constructor(e,t){let r=!(arguments.length>2&&void 0!==arguments[2])||arguments[2];super(e,t,P.t9,r),this.importAggregator()}}var R,I=i(2210),k=i(1214),H=i(2177),L={};try{R=localStorage.getItem("__nr_flags").split(","),console&&"function"==typeof console.log&&(L.console=!0,-1!==R.indexOf("dev")&&(L.dev=!0),-1!==R.indexOf("nr_dev")&&(L.nrDev=!0))}catch(e){}function z(e){try{L.console&&z(e)}catch(e){}}L.nrDev&&H.ee.on("internal-error",(function(e){z(e.stack)})),L.dev&&H.ee.on("fn-err",(function(e,t,r){z(r.stack)})),L.dev&&(z("NR AGENT IN DEVELOPMENT MODE"),z("flags: "+(0,b.D)(L,(function(e,t){return e})).join(", ")));var M=i(6660);class B extends h{static featureName=M.t;constructor(r,n){let i=!(arguments.length>2&&void 0!==arguments[2])||arguments[2];super(r,n,M.t,i),this.skipNext=0;try{this.removeOnAbort=new AbortController}catch(e){}const o=this;o.ee.on("fn-start",(function(e,t,r){o.abortHandler&&(o.skipNext+=1)})),o.ee.on("fn-err",(function(t,r,n){o.abortHandler&&!n[M.A]&&((0,I.X)(n,M.A,(function(){return!0})),this.thrown=!0,(0,s.p)("err",[n,(0,p.z)()],void 0,e.D.jserrors,o.ee))})),o.ee.on("fn-end",(function(){o.abortHandler&&!this.thrown&&o.skipNext>0&&(o.skipNext-=1)})),o.ee.on("internal-error",(function(t){(0,s.p)("ierr",[t,(0,p.z)(),!0],void 0,e.D.jserrors,o.ee)})),this.origOnerror=c._A.onerror,c._A.onerror=this.onerrorHandler.bind(this),c._A.addEventListener("unhandledrejection",(t=>{const r=function(e){let t="Unhandled Promise Rejection: ";if(e instanceof Error)try{return e.message=t+e.message,e}catch(t){return e}if(void 0===e)return new Error(t);try{return new Error(t+(0,D.P)(e))}catch(e){return new Error(t)}}(t.reason);(0,s.p)("err",[r,(0,p.z)(),!1,{unhandledPromiseRejection:1}],void 0,e.D.jserrors,this.ee)}),(0,O.m$)(!1,this.removeOnAbort?.signal)),(0,k.gy)(this.ee),(0,k.BV)(this.ee),(0,k.em)(this.ee),(0,t.OP)(r).xhrWrappable&&(0,k.Kf)(this.ee),this.abortHandler=this.#e,this.importAggregator()}#e(){this.removeOnAbort?.abort(),this.abortHandler=void 0}onerrorHandler(t,r,n,i,o){"function"==typeof this.origOnerror&&this.origOnerror(...arguments);try{this.skipNext?this.skipNext-=1:(0,s.p)("err",[o||new F(t,r,n),(0,p.z)()],void 0,e.D.jserrors,this.ee)}catch(t){try{(0,s.p)("ierr",[t,(0,p.z)(),!0],void 0,e.D.jserrors,this.ee)}catch(e){}}return!1}}function F(e,t,r){this.message=e||"Uncaught error with no additional information",this.sourceURL=t,this.line=r}let U=1;const q="nr@id";function G(e){const t=typeof e;return!e||"object"!==t&&"function"!==t?-1:e===c._A?0:(0,I.X)(e,q,(function(){return U++}))}function V(e){if("string"==typeof e&&e.length)return e.length;if("object"==typeof e){if("undefined"!=typeof ArrayBuffer&&e instanceof ArrayBuffer&&e.byteLength)return e.byteLength;if("undefined"!=typeof Blob&&e instanceof Blob&&e.size)return e.size;if(!("undefined"!=typeof FormData&&e instanceof FormData))try{return(0,D.P)(e).length}catch(e){return}}}var X=i(7243);class W{constructor(e){this.agentIdentifier=e,this.generateTracePayload=this.generateTracePayload.bind(this),this.shouldGenerateTrace=this.shouldGenerateTrace.bind(this)}generateTracePayload(e){if(!this.shouldGenerateTrace(e))return null;var r=(0,t.DL)(this.agentIdentifier);if(!r)return null;var n=(r.accountID||"").toString()||null,i=(r.agentID||"").toString()||null,o=(r.trustKey||"").toString()||null;if(!n||!i)return null;var a=(0,_.M)(),s=(0,_.Ht)(),c=Date.now(),u={spanId:a,traceId:s,timestamp:c};return(e.sameOrigin||this.isAllowedOrigin(e)&&this.useTraceContextHeadersForCors())&&(u.traceContextParentHeader=this.generateTraceContextParentHeader(a,s),u.traceContextStateHeader=this.generateTraceContextStateHeader(a,c,n,i,o)),(e.sameOrigin&&!this.excludeNewrelicHeader()||!e.sameOrigin&&this.isAllowedOrigin(e)&&this.useNewrelicHeaderForCors())&&(u.newrelicHeader=this.generateTraceHeader(a,s,c,n,i,o)),u}generateTraceContextParentHeader(e,t){return"00-"+t+"-"+e+"-01"}generateTraceContextStateHeader(e,t,r,n,i){return i+"@nr=0-1-"+r+"-"+n+"-"+e+"----"+t}generateTraceHeader(e,t,r,n,i,o){if(!("function"==typeof c._A?.btoa))return null;var a={v:[0,1],d:{ty:"Browser",ac:n,ap:i,id:e,tr:t,ti:r}};return o&&n!==o&&(a.d.tk=o),btoa((0,D.P)(a))}shouldGenerateTrace(e){return this.isDtEnabled()&&this.isAllowedOrigin(e)}isAllowedOrigin(e){var r=!1,n={};if((0,t.Mt)(this.agentIdentifier,"distributed_tracing")&&(n=(0,t.P_)(this.agentIdentifier).distributed_tracing),e.sameOrigin)r=!0;else if(n.allowed_origins instanceof Array)for(var i=0;i 2&&void 0!==arguments[2])||arguments[2];super(r,n,Z.t,i),(0,t.OP)(r).xhrWrappable&&(this.dt=new W(r),this.handler=(e,t,r,n)=>(0,s.p)(e,t,r,n,this.ee),(0,k.u5)(this.ee),(0,k.Kf)(this.ee),function(r,n,i,o){function a(e){var t=this;t.totalCbs=0,t.called=0,t.cbTime=0,t.end=E,t.ended=!1,t.xhrGuids={},t.lastSize=null,t.loadCaptureCalled=!1,t.params=this.params||{},t.metrics=this.metrics||{},e.addEventListener("load",(function(r){_(t,e)}),(0,O.m$)(!1)),c.IF||e.addEventListener("progress",(function(e){t.lastSize=e.loaded}),(0,O.m$)(!1))}function s(e){this.params={method:e[0]},T(this,e[1]),this.metrics={}}function u(e,n){var i=(0,t.DL)(r);i.xpid&&this.sameOrigin&&n.setRequestHeader("X-NewRelic-ID",i.xpid);var a=o.generateTracePayload(this.parsedOrigin);if(a){var s=!1;a.newrelicHeader&&(n.setRequestHeader("newrelic",a.newrelicHeader),s=!0),a.traceContextParentHeader&&(n.setRequestHeader("traceparent",a.traceContextParentHeader),a.traceContextStateHeader&&n.setRequestHeader("tracestate",a.traceContextStateHeader),s=!0),s&&(this.dt=a)}}function d(e,t){var r=this.metrics,i=e[0],o=this;if(r&&i){var a=V(i);a&&(r.txSize=a)}this.startTime=(0,p.z)(),this.listener=function(e){try{"abort"!==e.type||o.loadCaptureCalled||(o.params.aborted=!0),("load"!==e.type||o.called===o.totalCbs&&(o.onloadCalled||"function"!=typeof t.onload)&&"function"==typeof o.end)&&o.end(t)}catch(e){try{n.emit("internal-error",[e])}catch(e){}}};for(var s=0;s 1?e[1]=i:e.push(i)}else e[0]&&e[0].headers&&s(e[0].headers,n)&&(this.dt=n);function s(e,t){var r=!1;return t.newrelicHeader&&(e.set("newrelic",t.newrelicHeader),r=!0),t.traceContextParentHeader&&(e.set("traceparent",t.traceContextParentHeader),t.traceContextStateHeader&&e.set("tracestate",t.traceContextStateHeader),r=!0),r}}function x(e,t){this.params={},this.metrics={},this.startTime=(0,p.z)(),this.dt=t,e.length>=1&&(this.target=e[0]),e.length>=2&&(this.opts=e[1]);var r,n=this.opts||{},i=this.target;"string"==typeof i?r=i:"object"==typeof i&&i instanceof Y?r=i.url:c._A?.URL&&"object"==typeof i&&i instanceof URL&&(r=i.href),T(this,r);var o=(""+(i&&i instanceof Y&&i.method||n.method||"GET")).toUpperCase();this.params.method=o,this.txSize=V(n.body)||0}function A(t,r){var n;this.endTime=(0,p.z)(),this.params||(this.params={}),this.params.status=r?r.status:0,"string"==typeof this.rxSize&&this.rxSize.length>0&&(n=+this.rxSize);var o={txSize:this.txSize,rxSize:n,duration:(0,p.z)()-this.startTime};i("xhr",[this.params,o,this.startTime,this.endTime,"fetch"],this,e.D.ajax)}function E(t){var r=this.params,n=this.metrics;if(!this.ended){this.ended=!0;for(var o=0;o 2&&void 0!==arguments[2])||arguments[2];super(e,t,we.t,r),this.importAggregator()}}new class{constructor(e){let t=arguments.length>1&&void 0!==arguments[1]?arguments[1]:(0,_.ky)(16);c._A?(this.agentIdentifier=t,this.sharedAggregator=new y({agentIdentifier:this.agentIdentifier}),this.features={},this.desiredFeatures=new Set(e.features||[]),this.desiredFeatures.add(m),Object.assign(this,(0,a.j)(this.agentIdentifier,e,e.loaderType||"agent")),this.start()):(0,l.Z)("Failed to initial the agent. Could not determine the runtime environment.")}get config(){return{info:(0,t.C5)(this.agentIdentifier),init:(0,t.P_)(this.agentIdentifier),loader_config:(0,t.DL)(this.agentIdentifier),runtime:(0,t.OP)(this.agentIdentifier)}}start(){const t="features";try{const r=n(this.agentIdentifier),i=[...this.desiredFeatures];i.sort(((t,r)=>e.p[t.featureName]-e.p[r.featureName])),i.forEach((t=>{if(r[t.featureName]||t.featureName===e.D.pageViewEvent){const n=function(t){switch(t){case e.D.ajax:return[e.D.jserrors];case e.D.sessionTrace:return[e.D.ajax,e.D.pageViewEvent];case e.D.sessionReplay:return[e.D.sessionTrace];case e.D.pageViewTiming:return[e.D.pageViewEvent];default:return[]}}(t.featureName);n.every((e=>r[e]))||(0,l.Z)("".concat(t.featureName," is enabled but one or more dependent features has been disabled (").concat((0,D.P)(n),"). This may cause unintended consequences or missing data...")),this.features[t.featureName]=new t(this.agentIdentifier,this.sharedAggregator)}})),(0,T.Qy)(this.agentIdentifier,this.features,t)}catch(e){(0,l.Z)("Failed to initialize all enabled instrument classes (agent aborted) -",e);for(const e in this.features)this.features[e].abortHandler?.();const r=(0,T.fP)();return delete r.initializedAgents[this.agentIdentifier]?.api,delete r.initializedAgents[this.agentIdentifier]?.[t],delete this.sharedAggregator,r.ee?.abort(),delete r.ee?.get(this.agentIdentifier),!1}}}({features:[J,m,S,class extends h{static featureName=oe;constructor(t,r){if(super(t,r,oe,!(arguments.length>2&&void 0!==arguments[2])||arguments[2]),!c.il)return;const n=this.ee;let i;(0,k.QU)(n),this.eventsEE=(0,k.em)(n),this.eventsEE.on(se,(function(e,t){this.bstStart=(0,p.z)()})),this.eventsEE.on(ae,(function(t,r){(0,s.p)("bst",[t[0],r,this.bstStart,(0,p.z)()],void 0,e.D.sessionTrace,n)})),n.on(ce+ne,(function(e){this.time=(0,p.z)(),this.startPath=location.pathname+location.hash})),n.on(ce+ie,(function(t){(0,s.p)("bstHist",[location.pathname+location.hash,this.startPath,this.time],void 0,e.D.sessionTrace,n)}));try{i=new PerformanceObserver((t=>{const r=t.getEntries();(0,s.p)(te,[r],void 0,e.D.sessionTrace,n)})),i.observe({type:re,buffered:!0})}catch(e){}this.importAggregator({resourceObserver:i})}},C,xe,B,class extends h{static featureName=de;constructor(e,r){if(super(e,r,de,!(arguments.length>2&&void 0!==arguments[2])||arguments[2]),!c.il)return;if(!(0,t.OP)(e).xhrWrappable)return;try{this.removeOnAbort=new AbortController}catch(e){}let n,i=0;const o=this.ee.get("tracer"),a=(0,k._L)(this.ee),s=(0,k.Lg)(this.ee),u=(0,k.BV)(this.ee),d=(0,k.Kf)(this.ee),f=this.ee.get("events"),l=(0,k.u5)(this.ee),h=(0,k.QU)(this.ee),g=(0,k.Gm)(this.ee);function m(e,t){h.emit("newURL",[""+window.location,t])}function v(){i++,n=window.location.hash,this[ve]=(0,p.z)()}function b(){i--,window.location.hash!==n&&m(0,!0);var e=(0,p.z)();this[pe]=~~this[pe]+e-this[ve],this[ye]=e}function y(e,t){e.on(t,(function(){this[t]=(0,p.z)()}))}this.ee.on(ve,v),s.on(be,v),a.on(be,v),this.ee.on(ye,b),s.on(ge,b),a.on(ge,b),this.ee.buffer([ve,ye,"xhr-resolved"],this.featureName),f.buffer([ve],this.featureName),u.buffer(["setTimeout"+le,"clearTimeout"+fe,ve],this.featureName),d.buffer([ve,"new-xhr","send-xhr"+fe],this.featureName),l.buffer([me+fe,me+"-done",me+he+fe,me+he+le],this.featureName),h.buffer(["newURL"],this.featureName),g.buffer([ve],this.featureName),s.buffer(["propagate",be,ge,"executor-err","resolve"+fe],this.featureName),o.buffer([ve,"no-"+ve],this.featureName),a.buffer(["new-jsonp","cb-start","jsonp-error","jsonp-end"],this.featureName),y(l,me+fe),y(l,me+"-done"),y(a,"new-jsonp"),y(a,"jsonp-end"),y(a,"cb-start"),h.on("pushState-end",m),h.on("replaceState-end",m),window.addEventListener("hashchange",m,(0,O.m$)(!0,this.removeOnAbort?.signal)),window.addEventListener("load",m,(0,O.m$)(!0,this.removeOnAbort?.signal)),window.addEventListener("popstate",(function(){m(0,i>1)}),(0,O.m$)(!0,this.removeOnAbort?.signal)),this.abortHandler=this.#e,this.importAggregator()}#e(){this.removeOnAbort?.abort(),this.abortHandler=void 0}}],loaderType:"spa"})})(),window.NRBA=o})(); window.jQuery || document.write(' ') CKEDITOR_BASEPATH='https://f1000research.com/js/vendor/ckeditor/' window.reactTheme = 'research'; window.MathJax = { CommonHTML: { linebreaks: { automatic: true } }, 'HTML-CSS': { linebreaks: { automatic: true } }, SVG: { linebreaks: { automatic: true } }, AuthorInit: function() { MathJax.Hub.Register.MessageHook('End Process', function () { let timeout = false; // holder for timeout id const delay = 250; // delay after event is "complete" to run callback const reflowMath = function() { const dispFormulas = document.querySelectorAll('.disp-formula.panel'); if (!dispFormulas) { return; } for (const dispFormula of dispFormulas) { const child = dispFormula.querySelector('.MathJax_Preview').nextSibling.firstChild; const isMultiline = MathJax.Hub.getAllJax(dispFormula)[0].root.isMultiline; if (dispFormula.offsetWidth < child.offsetWidth || isMultiline) { MathJax.Hub.Queue(['Rerender', MathJax.Hub, dispFormula]); } } }; window.addEventListener('resize', function() { clearTimeout(timeout); // clear the timeout timeout = setTimeout(reflowMath, delay); // start timing for event "completion" }); }); }, }; if (window.location.hash == '#_=_'){ window.location = window.location.href.split('#')[0] } !function(f,b,e,v,n,t,s){if(f.fbq)return;n=f.fbq=function() {n.callMethod? n.callMethod.apply(n,arguments):n.queue.push(arguments)} ;if(!f._fbq)f._fbq=n; n.push=n;n.loaded=!0;n.version='2.0';n.queue=[];t=b.createElement(e);t.async=!0; t.src=v;s=b.getElementsByTagName(e)[0];s.parentNode.insertBefore(t,s)}(window, document,'script','https://connect.facebook.net/en_US/fbevents.js'); fbq('init', '1641728616063202'); fbq('track', "PixelInitialized", {}); (function(h,o,t,j,a,r){ h.hj=h.hj||function(){(h.hj.q=h.hj.q||[]).push(arguments)}; h._hjSettings={hjid:2318163,hjsv:6}; a=o.getElementsByTagName('head')[0]; r=o.createElement('script');r.async=1; r.src=t+h._hjSettings.hjid+j+h._hjSettings.hjsv; a.appendChild(r); })(window,document,'https://static.hotjar.com/c/hotjar-','.js?sv='); search file_upload Submit your research search menu close search Browse Gateways & Collections How to Publish Submit your Research My Submissions Article Guidelines Article Guidelines (New Versions) Open Data, Software and Code Guidelines Open Data and Accessible Source Materials Guidelines (HSS) Open Data, Software and Code Guidelines (PSE) Prepublication Checks Production Process Posters and Slides Guidelines Document Guidelines Article Processing Charges Peer Review Finding Article Reviewers About How it Works For Reviewers Our Advisors Policies Glossary FAQs For Developers Newsroom Contact My Research Submissions Content and Tracking Alerts My Details Sign In file_upload Submit your research { "@context": "https://schema.org", "@type": "ScholarlyArticle", "mainEntityOfPage": { "@type": "WebPage", "@id": "https://f1000research.com/articles/8-1032" }, "headline": "One-year test-retest reliability of ten vision tests in Canadian athletes", "datePublished": "2019-07-09T16:32:32", "dateModified": "2020-09-09T12:18:43", "author": [ { "@type": "Person", "name": "Mehdi Aloosh" }, { "@type": "Person", "name": "Suzanne Leclerc" }, { "@type": "Person", "name": "Stephanie Long" }, { "@type": "Person", "name": "Guowei Zhong" }, { "@type": "Person", "name": "James M. Brophy" }, { "@type": "Person", "name": "Tibor Schuster" }, { "@type": "Person", "name": "Russell Steele" }, { "@type": "Person", "name": "Ian Shrier" } ], "publisher": { "@type": "Organization", "name": "F1000Research", "logo": { "@type": "ImageObject", "url": "https://f1000research.com/img/AMP/F1000Research_image.png", "height": 480, "width": 60 } }, "image": { "@type": "ImageObject", "url": "https://f1000research.com/img/AMP/F1000Research_image.png", "height": 1200, "width": 150 }, "description": "Background: Vision tests are used in concussion management and baseline testing. Concussions, however, often occur months after baseline testing and reliability studies generally examine intervals limited to days or one week. Our objective was to determine the one-year test-retest reliability of these tests. Methods: We assessed one-year test-retest reliability of ten vision tests in elite Canadian athletes followed by the Institut National du Sport du Quebec. We included athletes who completed two baseline (preseason) annual evaluations by one clinician within 365±30 days. We excluded athletes with any concussion or vision training in between the annual evaluations or presented with any factor that is believed to affect the tests (e.g. migraines). Data were collected from clinical charts. We evaluated test-retest reliability using Intraclass Correlation Coefficient (ICC) and 95% limits of agreement (LoA). Results: We examined nine female and seven male athletes with a mean age of 22.7 (SD 4.5) years. Among the vision tests, we observed excellent test-retest reliability in Positive Fusional Vergence at 30cm (ICC=0.93) but this dropped to 0.53 when an outlier was excluded in a sensitivity analysis. There was good to moderate reliability in Negative Fusional Vergence at 30cm (ICC=0.78), Phoria at 30cm (ICC=0.68), Near Point of Convergence break (ICC=0.65) and Saccades (ICC=0.61). The ICC for Positive Fusional Vergence at 3m (ICC=0.56) also decreased to 0.45 after removing two outliers. We found poor reliability in Near Point of Convergence (ICC=0.47), Gross Stereoscopic Acuity (ICC=0.03) and Negative Fusional Vergence at 3m (ICC=0.0). ICC for Phoria at 3m was not appropriate because scores were identical in 14/16 athletes. 95% LoA of the majority of tests were ±40% to ±90%. Conclusions: Five tests had good to moderate one-year test-retest reliability. The remaining tests had poor reliability. The tests would therefore be useful only if concussion has a moderate-large effect on scores." } { "@context": "http://schema.org", "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": "1", "item": { "@id": "https://f1000research.com/", "name": "Home" } }, { "@type": "ListItem", "position": "2", "item": { "@id": "https://f1000research.com/browse/articles", "name": "Browse" } }, { "@type": "ListItem", "position": "3", "item": { "@id": "https://f1000research.com/articles/8-1032/v5", "name": "One-year test-retest reliability of ten vision tests in Canadian athletes" } } ] } Home Browse One-year test-retest reliability of ten vision tests in Canadian athletes ALL Metrics - Views Downloads Get PDF Get XML Cite How to cite this article Aloosh M, Leclerc S, Long S et al. One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.12688/f1000research.19587.5 ) NOTE: If applicable, it is important to ensure the information in square brackets after the title is included in all citations of this article. Close Copy Citation Details Export Export Citation Sciwheel EndNote Ref. Manager Bibtex ProCite Sente EXPORT Select a format first Track Share ▬ ✚ Research Article Revised One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] Mehdi Aloosh https://orcid.org/0000-0001-6763-1067 1,2 , Suzanne Leclerc 3 , Stephanie Long 4 , [...] Guowei Zhong 4 , James M. Brophy 5 , Tibor Schuster 4 , Russell Steele 6 , Ian Shrier 4,7 Mehdi Aloosh https://orcid.org/0000-0001-6763-1067 1,2 , Suzanne Leclerc 3 , [...] Stephanie Long 4 , Guowei Zhong 4 , James M. Brophy 5 , Tibor Schuster 4 , Russell Steele 6 , Ian Shrier 4,7 PUBLISHED 09 Sep 2020 Author details Author details 1 Department of Epidemiology, Biostatistics and Occupational Health, McGill University, Montreal, Canada 2 Department of Health Research Methods, Evidence, and Impact, Michael G. DeGroote School of Medicine, McMaster University, Hamilton, Canada 3 Institut National du Sport du Quebec, Montreal, Canada 4 Department of Family Medicine, McGill University, Montreal, Canada 5 Faculty of Medicine, McGill University, Montreal, Canada 6 Department of Mathematics and Statistics, McGill University, Montreal, Canada 7 Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada Mehdi Aloosh Roles: Writing – Original Draft Preparation, Writing – Review & Editing Suzanne Leclerc Roles: Conceptualization, Funding Acquisition, Resources, Writing – Review & Editing Stephanie Long Roles: Data Curation, Writing – Review & Editing Guowei Zhong Roles: Data Curation, Writing – Review & Editing James M. Brophy Roles: Validation, Writing – Review & Editing Tibor Schuster Roles: Formal Analysis Russell Steele Roles: Formal Analysis, Writing – Review & Editing Ian Shrier Roles: Conceptualization, Formal Analysis, Funding Acquisition, Investigation, Resources, Supervision, Writing – Review & Editing OPEN PEER REVIEW DETAILS REVIEWER STATUS Abstract Background : Vision tests are used in concussion management and baseline testing. Concussions, however, often occur months after baseline testing and reliability studies generally examine intervals limited to days or one week. Our objective was to determine the one-year test-retest reliability of these tests. Methods : We assessed one-year test-retest reliability of ten vision tests in elite Canadian athletes followed by the Institut National du Sport du Quebec. We included athletes who completed two baseline (preseason) annual evaluations by one clinician within 365±30 days. We excluded athletes with any concussion or vision training in between the annual evaluations or presented with any factor that is believed to affect the tests (e.g. migraines). Data were collected from clinical charts. We evaluated test-retest reliability using Intraclass Correlation Coefficient (ICC) and 95% limits of agreement (LoA). Results: We examined nine female and seven male athletes with a mean age of 22.7 (SD 4.5) years. Among the vision tests, we observed excellent test-retest reliability in Positive Fusional Vergence at 30cm (ICC=0.93) but this dropped to 0.53 when an outlier was excluded in a sensitivity analysis. There was good to moderate reliability in Negative Fusional Vergence at 30cm (ICC=0.78), Phoria at 30cm (ICC=0.68), Near Point of Convergence break (ICC=0.65) and Saccades (ICC=0.61). The ICC for Positive Fusional Vergence at 3m (ICC=0.56) also decreased to 0.45 after removing two outliers. We found poor reliability in Near Point of Convergence (ICC=0.47), Gross Stereoscopic Acuity (ICC=0.03) and Negative Fusional Vergence at 3m (ICC=0.0). ICC for Phoria at 3m was not appropriate because scores were identical in 14/16 athletes. 95% LoA of the majority of tests were ±40% to ±90%. Conclusions: Five tests had good to moderate one-year test-retest reliability. The remaining tests had poor reliability. The tests would therefore be useful only if concussion has a moderate-large effect on scores. READ ALL READ LESS Keywords concussion, vision tests, binocular, saccades, reliability Corresponding Author(s) Ian Shrier ( [email protected] ) Close Corresponding author: Ian Shrier Competing interests: No competing interests were disclosed. Grant information: This project was funded by government sources (MITACS [IT08159] and MEDTEQ [G245120], a non-profit organization (Institut National du Sport du Quebec) and private industry (APEXk Inc, and Varitron Inc.) The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Copyright: © 2020 Aloosh M et al . This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. How to cite: Aloosh M, Leclerc S, Long S et al. One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.12688/f1000research.19587.5 ) First published: 09 Jul 2019, 8 :1032 ( https://doi.org/10.12688/f1000research.19587.1 ) Latest published: 09 Sep 2020, 8 :1032 ( https://doi.org/10.12688/f1000research.19587.5 ) Revised Amendments from Version 4 The reviewers insisted that we consider our sensitivity analyses excluding outliers as the primary analysis. Although we do not believe this is the optimal approach, we made this change in the current version. There are some other small edits about some values from this athlete population being outside the normative data in the general population that we obtained from the literature. The reviewers insisted that we consider our sensitivity analyses excluding outliers as the primary analysis. Although we do not believe this is the optimal approach, we made this change in the current version. There are some other small edits about some values from this athlete population being outside the normative data in the general population that we obtained from the literature. See the authors' detailed response to the review by M Nadir Haider See the authors' detailed response to the review by Dillon Richards and James P Dickey READ REVIEWER RESPONSES Introduction Concussion, a form of mild traumatic brain injury is a growing public health concern 1 . Estimates suggest up to 3.8 million sport-related concussions occur annually in the United States, with 50% going unreported 2 . United States emergency department visits for sports-related traumatic brain injuries have increased 60% over 2001–2009 3 . Concussions can be associated with headaches, dizziness, visual disturbances, and other symptoms that can negatively affect performance in sport, school, and work and negatively impact quality of life 2 , 4 , 5 . Diagnosis of concussion and decisions to return-to-play are based on symptoms, signs, physical examination and special tests 6 . Previous research has shown an association between concussion and eye movement 1 . Concussion may therefore affect multiple aspects of vision, including saccades, pursuit, convergence, accommodation, and vestibulo-ocular reflex 7 . Some studies reported 50% to 90% incidence of visual symptoms, such as blurred vision and diplopia in individuals with concussion 8 . Therefore, vision testing may be helpful in the assessment and management of patients with concussion. Each vision test measures a function that is linked to a particular brain structure or pathway. Vision tests are noninvasive tests with rapid administration and scoring. Understanding test variability, independent of changes in pathology or recovery (i.e. reliability), is required to assess their clinical utility. However, only a limited number of reliability studies have assessed binocular vision tests and saccades 9 – 20 . In addition, these reliability studies measured a specific aspect of the vision. These studies are not uniform in their method and they are diverse in their population. Previous investigations of the test-retest reliability of these vision tests have used short test-retest time intervals ranging from 0 to approximately 57 days 9 – 20 , except for one test of saccades 21 . For test-retest reliability to be useful in clinical management (e.g. return-to-play), the time intervals must reflect the time frame in which they would be used 22 . The previous studies have provided information on the usefulness of these tests when following improvement or deterioration of patients over short periods of time. However, concussions usually occur several months and up to one year after annual baseline testing, and not as 0 days to 57 days as in the previous studies. Therefore, we examined one-year test-retest reliability of ten vision tests in Canadian athletes over one year period of time. Methods Participants The study population included athletes over 16 years of age followed by the Institut National du Sport du Quebec (INSQ) in Canada from 2015–2018. Many of these athletes had a yearly examination done by a sports medicine physician and vision tests done by a clinician trained in orthoptic testing. We only included athletes who had completed two baseline (preseason) annual evaluations within a 365-day (± 30 days) time period. We excluded athletes who suffered a concussion in between annual evaluations or had received preventive orthoptic training between the baseline measures. We also excluded athletes with a history of strabismus or treated strabismus, or were medically treated for depression, anxiety or psychiatric conditions that may affect binocular vision and saccades. Data were collected from electronic medical charts of one clinician trained in orthoptic measures and one sports medicine physician. Measures At the beginning of each season, athletes underwent baseline testing of ten vision tests by a single orthoptic-trained clinician (industry partner). The vision tests were Gross Stereoscopic Acuity, Near Point of Convergence (NPC), Near Point of Convergence break (NPCb), near (30cm) and far (3m) Positive Fusional Vergence, near (30cm) and far (3m) Negative Fusional Vergence, near (30cm) and far (3m) Phoria, and Saccades. A detailed description of each test including the procedures of each test and the theoretical range of scores is provided in Table 1 . We will briefly describe each vision test here. We used a horizontal prism bar with the base-out for Positive Fusional Vergence and base-in for Negative Fusional Vergence, at both 30cm and 3m 10 . Phoria was measured at 30cm and 3m using the prism and alternate cover test using the procedures described by the Pediatric Eye Disease Investigator Group 23 . To perform NPC and NPCb, we followed the Maples et al ., protocol 13 . We measured Gross Stereoscopic Acuity with the Randot Stereotest (Stereo Optical Co., Inc., Chicago, IL) according to the manufacturer’s instructions 24 . Evaluation of Saccades was done using the test procedures developed by the orthoptic-trained clinician. Participants assumed a tandem stance an arm’s length away from a screen attempting to fixate on appearing and disappearing lights on the screen, while trying to keep their head still. Light flashes appeared at a rate of 100 per minute for two minutes. This test was scored by the clinician based on quality (bad, medium, good), synchronization (bad, medium, good), and saccadic corrections (many, few, none). These three components were then combined into an overall percentage saccade score, based on an unpublished proprietary algorithm developed by the clinician who performed the testing. Table 1. Detailed description of the ten vision tests. Positive Fusional Vergence This test examines how well a participant can adapt to challenges in fixating light on their retina at near distance (30cm) and far distance (3m), measured in prism diopters. The seated participant fixates on a fixed target at the appropriate distance. The clinician begins by using the weakest prism strength (base-out) which forces the participant to converge their eyes to maintain fixation. The strength of the prism is increased until the participant can no longer maintain a single image. The score of each test (30cm and 3m) is the strength of the prism in which the participant maintained binocular vision, with higher scores representing better function. The range of normative data for Positive Fusional Vergence at near fixation is 35 to 40 prism diopters, and the range at far fixation is 16 to 20 prism diopters 25 – 27 . Negative Fusional Vergence This is the same test as Positive Fusional Vergence except the horizontal prism bar is positioned base-in, forcing the participant to diverge their eyes to maintain fixation on a fixed object positioned at near (30cm) and far (3m), measured in prism diopters. The clinician incrementally increases the strength of the prism until the participant is no longer able to maintain a single image. The score of each test is the strength of the prism in which the participant maintained binocular vision, with higher scores representing better function. The range of normative data for Negative Fusional Vergence at near fixation is 12 to 16 prism diopters, and the range at far fixation is 6 to 8 prism diopters 25 – 27 . Phoria We evaluated the natural deviation of the eyes (heterophoria), in prism diopters, with the prism and alternate cover test using a target placed at (1) 3m from the participant (far vision), and (2) 30cm from the participant (near vision). While the seated participant was fixating on the target, the clinician covered and uncovered each of the participant’s eyes to trigger movements while using a prism bar (base-out if the eye moves outward, base-in if the eye moves inward) to cancel these movements. The prism power was progressively increased until no shift in the eyes was seen. The score of the test was the rating of the prism that canceled the eye movements, with lower scores representing less Phoria. We were unable to find normative data for this test. Near Point of Convergence (NPC) NPC assesses the ability to symmetrically converge, and is sometimes referred to as “motor punctum proximum” 26 , in cm. The seated participant fixates on a near target 30cm away. The target is gradually moved towards their eyes as they attempt to maintain fixation. NPC is reached when one or both eyes can no longer maintain fixation on the target, which is identified as when one eye diverges outwards. The score of the test is the distance (cm) between the bridge of the nose and the distance of the target at the closest point at which the individual could maintain balanced oculomotor synergy between both eyes. Lower scores indicate better NPC. Normative data in older textbooks report average NPC values for healthy adults between 6 to 8 cm 28 , but a more recent study suggested 5 cm should be considered the upper limit of normal values 29 . Near Point of Convergence break (NPCb) This test is conducted using the same methods as NPC, but the test ends when the participant has double vision due to the inability of the eyes to converge. The score of the test is the distance between the bridge of the nose and the point (in cm) where double vision occurs, where a lower score indicates better NPCb. Normative data for elementary school children with normal vision suggested a mean of 3.3 cm, with a range of 1.0 to 13.7 cm 30 ; however, data on adults with normal vision suggest a breakpoint of approximately 5.0 to 7.5 cm 31 . Gross Stereoscopic Acuity We tested the ability to perceive depth with the Randot® Stereotest (Stereo Optical Co., Inc., Chicago, IL), in arc seconds. Seated participants wearing polarized glasses were asked to hold the testing booklet 16 inches from their face. Participants were then presented images formed of dots that are displaced in relation to each other. The test steadily increased in difficulty by reducing the level of disparity between dots, beginning at 400 arc seconds (lowest possible score) and ending at 20 arc seconds (highest possible score). A participant’s score was the arc seconds corresponding to the smallest disparity at which the participant identified the raised (i.e. stereoscopic) image. Normative data suggest the average score for an adult is 40 arc seconds 32 , 33 . Saccades This test examines the eye’s ability to perform saccadic movements, which are rapid eye movements that abruptly alter the point of fixation. In our clinician’s version of this test, participants assume a tandem stance (heel-to-toe with dominant foot in the back) standing an arm’s length away from the screen. Lights appear and disappear in different locations on the screen at a rate of 100 flashes per minute, for a total of two minutes. The participant is instructed to keep their head still and only move their eyes to fixate on the appearing lights. The clinician observes the eyes for quality and synchronization (rated: bad, medium, good) and saccadic correction (rated: many corrections, few corrections, no corrections). The three sub- scores were combined into an overall percentage score according to a proprietary algorithm developed by the clinician (industry partner) who performed the testing. There are no normative data for this version of the test because the score is based on a proprietary algorithm. Analysis We report the mean (SD) for continuous variables at baseline. We evaluated test-retest reliability using Intraclass Correlation Coefficient (ICC) 34 and 95% limits of agreement (LoA) 35 . We considered ICC of ≤0.5 as poor, 0.51–0.74 as moderate, 0.75–0.89 as good, and ≥0.90 as excellent reliability 36 . We report the LoA in the raw units of the scale used by clinicians. To compare LoA across tests, we also standardized the scores and reported them as percent differences, [(T1- T2)/ mean(T1&T2)]*100 35 , 37 . Additionally, we summarized LoA graphically with Bland-Altman plots for each vision test using the standardized score for the y-axis to provide an overview of all vision tests. The raw scale measures are provided in parentheses to provide clinicians with information for individual patient assessment. Finally, we conducted a sensitivity analysis for the vision tests by excluding outliers that may have augmented the ICC results. We defined an outlier as a data point that was 1.5 interquartile ranges below the first quartile or above the third quartile. Due to the limited sample size (n=16) and to avoid being overly conservative in our evaluation, we followed the practical solution for addressing multiple testing proposed by Saville, the unrestricted least significant difference procedure (or multiple t-test) 38 . Formal multiplicity correction of confidence levels was not performed but we thoroughly reported all statistical assessments enabling an informal type-I error assessment by the reader. The data were analyzed using R statistical software 3.4.3 39 . This study was approved by the McGill University Faculty of Medicine Institutional Review Board. Results Of the 199 athletes measured for the vision tests, only 16 individuals met our inclusion criteria ( Figure 1 ). There were nine female and seven male athletes with a mean age of 22.7 (4.5) years at the baseline (preseason) measurement. Participants were athletes of water polo (n=6) and short-track speed skating (n=10). A second measurement was conducted between 335 and 372 days (mean of 356.4 (17.3) days) after the initial baseline. Figure 1. Patient flow diagram. The range of scores observed for each vision test can be found in each of the reliability figures ( Figure 2 – Figure 4 ) 40 . Our analysis suggested one-year test-retest reliabilities ranging from poor to excellent among the ten vision tests. Including all the data, we observed excellent one-year test-retest reliability in Positive Fusional Vergence at 30cm with ICC of 0.93 ( Figure 2 ). In this test, 4 out of 16 pairs of measurements were identical after 1 year. The range of measurements was between 14 and 45 diopters with one outlier at 90 diopters. LoA of the test was ±41.9%. Given the very high ICC and the presence of an outlier that greatly increased the range of the values for the measure (known to increase ICC), we repeated the analysis excluding the outlier. This decreased the ICC from 0.93 to 0.53, and increased the LoA to ±43.5%. One of the reviewers for this paper has insisted that the analysis without the outlier be considered the primary analysis. Figure 2. Vision test with excellent one-year test-retest reliability. ( A ) Scatter plot of test-retest reliability for Positive Fusional Vergence at 30cm. Identity line represents perfect agreement between the test-retest values; ICC refers to the Intraclass correlation coefficient and 95%CI refers to the 95% Confidence Interval. “n (1,2,3,4)” refers to the number of participants represented by each dot when scores exactly overlapped. ( B ) Bland-Altman plot with the mean of the test-retest on the x-axis and the difference between test-retest on the y-axis. Solid line represents the bias and dotted lines represent the 95% LoA. The y-axis represents a standardized LoA using percentage difference on the plot to allow one to compare the different tests to each other. The LoA in the units of measure, which are familiar to clinicians, are provided in the parentheses. When the analysis was repeated excluding the outlier to the far right, the ICC decreased to 0.53 and the 95% LoA increased to 43.5%. Figure 3. Vision tests with good to moderate one-year test-retest reliability. ( A ) Scatter plot of test-retest for Negative Fusional Vergence at 30cm, Phoria at 30cm, Near Point of Convergence break (NPCb), Positive Fusional Vergence at 3m, and Saccades. ( B ) Bland-Altman plot related to each test. See Figure 2 for explanation of abbreviations and scales. When the analysis for Positive Fusional Vergence at 3m was repeated excluding the two outliers, the ICC decreased to 0.45 and the 95% LoA decreased to 41.4%. Figure 4. Vision tests with poor one-year test-retest reliability. ( A ) Scatter plots of test-retest for near point of convergence (NPC), Gross Stereoscopic Acuity, and Negative Fusional Vergence at 3m. ( B ) Bland-Altman plots related to each test. See Figure 2 for explanation of abbreviations and scales. Five tests showed good to moderate one-year test-retest reliability ( Figure 3 ), including Negative Fusional Vergence at 30cm (ICC=0.78, LoA=41.2%), Phoria at 30cm (ICC=0.68, LoA=119.2%), NPCb (ICC=0.65, LoA=49.4%), Positive Fusional Vergence at 3m (ICC=0.56, LoA=60.2%), and Saccades (ICC=0.61, LoA=24.3%). There were two outliers for Positive Fusional Vergence at 3m (one participant on both measures and one participant on only one measure). When we removed both of these outliers, the ICC dropped from 0.56 to 0.45 and the 95% LoA decreased from 60.2% to 41.4%. In both of these cases, the two scores from the outlier were quite different. Although one might anticipate that the ICC would increase by removing such outliers, the ICC actually decreased because the range of values for the measure decreased substantially. As above, one of the reviewers for this paper insisted that the analysis without the outliers be considered the primary analysis. Three of the remaining four tests showed poor one-year test-retest reliability ( Figure 4 ). These include NPC (ICC=0.47, LoA=73.9%), Gross Stereoscopic Acuity (ICC=0.03, LoA=92.5%) and Negative Fusional Vergence at 3m (ICC=0.0, LoA=48.4%). For Phoria at 3m, 14/16 athletes had identical scores on the two measures. In this context, the ICC and LoA were not appropriate measures of reliability and are not presented. Discussion We found that the one-year test-retest reliability for 10 vision tests in young elite athletes ranged from moderate to poor after accounting for outliers. The majority of the vision tests had standardized 95% LoA in the range of 40–90%, which indicates that repeated scores of an individual over time may vary by 40–90% of the mean score even without any actual change in vision function. There are a limited number of test–retest reliability studies on non-vision neurocognitive tests over a one year period in teenage athletes. For instance, the ICC for different components of Immediate Post-Concussion Assessment and Cognitive Testing (ImPACT), a computerized brain injury measurement tool, ranges from 0.50 to 0.82 41 . However, we could not find any research examining the stability of the vision tests over a one year period, in athlete or non-athlete populations except for one test of saccades that was very different from the test used in this study 21 . It is important that test-retest reliabilities fall within a range needed for clinical interpretation of concussion assessment and for discussion about return-to-play. In the context of comparing results after a concussion to annual baseline tests conducted in the pre-season, the time-frame for reliability comparisons should be up to one year 22 . Although there are no long-term reliability studies on the ten vision tests evaluated in this study, a number of studies have reported short term test-retest reliability of individual tests using various methods among various groups of individuals, including children and healthy adults 9 – 20 . Using NPC as a general example, one study reported excellent immediate test-retest reliability in concussed athletes (ages 9–24) (ICC = 0.95 to 0.98) 12 . A separate study using a 2–3 day test-retest protocol found the ICC = 0.65 for NPC in healthy individuals (calculated in Rouse et al. , 2002 16 for data from reference 15 ), and a third study reported one week test-retest ICC = 0.89 and 0.92 for NPCb in healthy school children 16 . We recently examined one-week test-retest reliability of the same ten vision tests with the same methods and same age-range as this current study in 20 young non-athletes. We found one-week test-retest reliability ranging from poor (ICC = 0.34) to good (ICC = 0.88), with five out of ten tests showing moderate reliability (ICCs = 0.54 to 0.69) 17 . This suggests that these vision tests can only be useful if a concussion has a moderate to large effect on scores. Overall, the ICCs in the current study were generally smaller than those reported in our one-week study, suggesting increased temporal variability. Unexpectedly, the 95% LoA for one-year test-retest was smaller or equal to the 95% LoA of the one-week test-retest for all vision tests except NPC (±73.9 vs. ±57.9) and Gross Stereoscopic Acuity (±92.5 vs. ±55). In addition, in both the one-week and one-year intervals, almost all individuals had the same value in Phoria 3m, which leads to uninformative LoA. In one-year test-retest, Positive Fusional Vergence showed excellent reliability at 30cm (ICC=0.93) and moderate at 3m (ICC=0.56), initially. Our results at 30cm were significantly better than those of another study examining test-retest reliability of Positive Fusional Vergence at 30cm in children (ICCs of 0.53–0.59) 16 . Perhaps more importantly, our results were also better than the one-week test-retest reliability conducted by the same clinician with the same methods in our previous prospective research study (ICC=0.54 and 0.49, respectively) 17 . It is difficult to understand how test-retest reliability over one year could be better than test-retest reliability over one week. When we explored the data further, we noticed one outlier that greatly increased the range of values for Positive Fusional Vergence at 30cm ( Figure 2 ) and Positive Fusional Vergence at 3m ( Figure 3 ). Increasing the range of values is known to increase the ICC. This is because ICC is based on the results of an analysis of variance which separates the error into variability between individuals (range of values along x or y axes) and variability within an individual. Therefore, if variability between persons increases, indicated by a larger range of values, ICC will increase. When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is similar to the value found for the one-week test-retest reliability (ICC=0.54); the LoA increased to 43.5%. When we removed the two outliers from Positive Fusional Vergence at 3m, the ICC decreased to 0.45 and LoA decreased to 41.4%. Note that the outliers for this measure had large differences between the two test scores, and removing such data points would normally be expected to increase the ICC ( Figure 3 ). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability. In addition to Positive Fusional Vergence, two other tests also had higher ICC at one year (Negative Fusional Vergence 30cm: 0.78 vs 0.66) and Saccades (0.61 vs 0.34) but there were no apparent outliers and the range of values were similar in the two studies. Aside from outliers, there are other theoretical reasons that might explain why ICC is better at one-year than at one-week. First, it is possible that the non-athletes in our one-week test-retest study had less motivation to perform well on the repeat tests. If true, their scores would be less than the motivated athletes performing during the one-year test-retest. Second, there is a potential learning effect in retest measurements that could affect results. A learning effect, however, is unlikely in our study because the athletes were tested only twice, with a one-year interval between tests. Third, the one-week study was a prospective research study where the clinician performing the test was blinded. Our current results are based on clinical charts where the clinician had access to the previous results which might artificially increase the reliability of the test. Fourth, the increased ICC could have occurred simply by chance because of sampling variation. Our measurements of Phoria at 30cm had moderate reliability for near (ICC=0.68) consistent with our one-week retest reliability study (ICC=0.69) 17 . Other studies in adults and children with strabismus 42 or esotropia 23 have not reported ICC. Therefore, comparing between studies is not possible. Moreover, our analytical methods differed slightly from those studies. We evaluated all angles of deviation together, and other authors analyzed smaller (2–20 Prism Diopter) or larger (>20 Prism Diopter) angles of strabismus separately because of different prism increments measured 42 . For Phoria at 3m, we found that the ICC and LoA were not appropriate measures of reliability because most of the population reported identical scores of zero for both measurements. One may consider that if we had a wider range of scores, ICC might provide meaningful information. One-year test-retest reliability of NPC and NPCb (0.47 and 0.65, respectively) were similar to the results in our one-week reliability study (0.54 and 0.64, respectively) 17 . Brozek et al . found a similar ICC of 0.65 for NPC in healthy adults (calculated in Rouse et al. , 2002 16 for data from Brozek et al. , 1948 15 ). However, Giffard et al. reported a one-week ICC = 0.84 in patients for NPC with neck pain 18 and Rouse et al. reported excellent one-week reliability for NPCb in school children (ICC=0.89 and 0.92 for two different examiners) 16 . The discrepancies in results are most likely due to differences in testing procedures. For instance, we used the Maples method 13 which is a non-accommodative test. Rouse et al. 16 used an accommodative target with Astron International Accommodative Rule and Giffard et al. 18 used the RAF rule 28 . Our one-year test-retest results for Gross Stereoscopic Acuity in young athletes showed poor reliability (ICC=0.03; 95% LoA= ±92.5%) even though our previous one-week test-retest results reported good reliability in non-athlete young adults (ICC=0.86; 95% LoA = ± 54%) 17 and another study using Titmus stereo fly and Frisby stereo tests in pre-school children revealed an excellent one-week reliability (ICC=1.0) 19 . In addition, another study reported that 82.0% of their participants had identical results at test and retest taken on the same day in 100 healthy adult and children 11 . With a one-year ICC of 0.03 and LoA of 92.5%, Gross Stereoscopic Acuity cannot be considered a reliable test to assess the vision function over one year, although it may still be appropriate for use in shorter time intervals, such as one week 11 , 17 , 19 . Finally, our clinician’s test of Saccades showed moderate reliability (ICC=0.61) with the smallest LoA (in percentage) of other tests, similar to the one-week study 17 . These results are similar to other findings in healthy adults over a two-month period (ICC=0.59) 20 . With a moderate reliability and the smallest LoA amongst the other vision tests, the results of the test of Saccades could be considered stable over a one year period assessing athletes. In this study, four vision tests (Negative Fusional Vergence at 30cm, Phoria at 30cm, Saccades and NPCb) had moderate one-year test-retest reliability. The one test with identical scores in 14/16 athletes was Phoria at 3m. Therefore we cannot comment on the reliability of this test. This level of reliability would be useful in conditions where the concussion leads to a moderate change in vision function. The remaining five vision tests, including Positive Fusional Vergence at 30cm and 3m, NPC, Negative Fusional Vergence at 3m, and Gross Stereoscopic Acuity may be useful to detect the effect of concussion with a large change on vision function. Further studies are therefore required to assess the effect of concussion on vision test scores of the five vision tests. If it can be shown that the concussion has moderate to large effect on the test scores then these vision tests may still be useful clinically. Strengths and limitations Several studies have previously evaluated the inter-rater reliability of some vision tests 23 , 42 . However, inter-rater reliability is less important in the context of clinical care when patients are followed by one clinician over time. Our study evaluated the test-retest reliability of the ten vision tests over an interval that allows for the normal variation over time expected in clinical practice between baseline measures and subsequent concussions. The ICC represents how much of variability in scores is due to differences between subjects. For instance, the ICC of 0.78 for near Negative Fusional Vergence at 30cm suggests that 78% of the variability in the measurements was due to differences between participants, and 22% was due to normal variations within the measurement. Furthermore, the 95% LoA for each test in our study provides the magnitude of the normal variation that can be expected with repeated measurements. Differences in test results between baseline and diagnosis of a concussion likely represent a true signal of a change in vision function within the patient if these differences are larger than the noise (LoA). In addition, we conducted sensitivity analysis to evaluate the effect of outliers. This analysis suggested that our initial ICC results may have been artificially high for two tests. (Positive Fusional Vergence at 30cm and 3m). Finally, the results of the test of Saccades in this study are based on the unpublished proprietary algorithm developed by the clinician. This limits its applicability for other clinicians. This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm). Some data in these athletes appear to be outside the normative range of data previously described for the general population. Because the data were obtained as part of clinical practice, the clinician had access to the results of the first test when conducting the repeat test one year later. The lack of blinding may result in higher agreement between the two tests compared to our blinded one-week research study. However, clinicians are not blinded during normal clinical practice, and therefore the results of this study would represent an expected level of agreement in that context, even if some of the agreement is due to bias. In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. Despite starting with athletes from many sports, only athletes from water polo and short-track speed skating met our eligibility criteria. It is unclear if subconcussion impacts affect neurological function in general 43 . If subconcussion impacts were common in these sports and affected vision testing, we should have seen a systematic decrease in vision capacity between the two tests; this was not observed. Further, if it were present, the effect would be considered part of the “noise” clinicians have to consider when comparing the results from post-concussion and baseline tests. With an effective sample size of 16, the anticipated precision of ICC estimates was +/- 0.25 and the study had 80% power to detect ICC values >= 0.6 and more than 90% power to detect ICC values >=0.7 i.e. rejection of the null hypothesis (Table 1a in 44 ). Note that a total of >60 individuals were required to exclude ICC values 0.7 (Table 2b in 44 ). Conclusion We found that five out of the ten vision tests (Negative Fusional Vergence at 30cm, Phoria at 30cm, NPCb, Positive Fusional Vergence at 30cm, and Saccades) had good to moderate one-year test-retest reliability. This level of reliability is useful in conditions which produce a moderate change in vision function. The remaining five vision tests may be useful in detecting large effects on vision function. If further studies suggest that the effect of concussion on test scores is moderate to large, these vision tests may still be useful clinically. Data availability Open Science Framework: Vision Tests in Concussion. https://doi.org/10.17605/OSF.IO/VB4W8 40 Data are available under the terms of the Creative Commons Attribution 4.0 International license (CC-BY 4.0). Demographic data are not available. With only 9 males and 7 females from our clinical source, any demographic information would immediately allow some participants to be identified and therefore this information cannot be shared in order to preserve participant confidentiality. Acknowledgments We would like to thank Isabel Pereira for her help throughout the course of this work. In addition, we would like to thank David Tinjust, from Apexk for examining the athletes. Faculty Opinions recommended References 1. Sussman ES, Ho AL, Pendharkar AV, et al. : Clinical evaluation of concussion: The evolving role of oculomotor assessments. Neurosurg Focus. 2016; 40 (4): E7. PubMed Abstract | Publisher Full Text 2. Langlois JA, Rutland-Brown W, Wald MM: The epidemiology and impact of traumatic brain injury: a brief overview. J Head Trauma Rehabil. 2006; 21 (5): 375–8. PubMed Abstract | Publisher Full Text 3. Centers for disease control and prevention: Nonfatal traumatic brain injuries related to sports and recreation activities among persons aged ≤19years--United States, 2001-2009. MMWR Morb Mortal Wkly Rep. 2011; 60 (39): 1337–42. PubMed Abstract 4. Dikmen S, Machamer J, Fann JR, et al. : Rates of symptom reporting following traumatic brain injury. J Int Neuropsychol Soc. 2010; 16 (3): 401–11. PubMed Abstract | Publisher Full Text 5. McCrory P, Meeuwisse WH, Aubry M, et al. : Consensus statement on concussion in sport: The 4th international conference on concussion in sport held in zurich, november 2012. Br J Sports Med. 2013; 47 (5): 250–8. PubMed Abstract | Publisher Full Text 6. McCrory P, Meeuwisse W, Dvořák J, et al. : Consensus statement on concussion in sport-the 5 th international conference on concussion in sport held in Berlin, October 2016. Br J Sports Med. 2017; 51 (11): 838–47. PubMed Abstract | Publisher Full Text 7. Ventura RE, Balcer LJ, Galetta SL: The neuro-ophthalmology of head trauma. Lancet Neurol. 2014; 13 (10): 1006–16. PubMed Abstract | Publisher Full Text 8. Talavage TM, Nauman EA, Breedlove EL, et al. : Functionally-detected cognitive impairment in high school football players without clinically-diagnosed concussion. J Neurotrauma. 2014; 31 (4): 327–38. PubMed Abstract | Publisher Full Text | Free Full Text 9. Oberlander TJ, Olson BL, Weidauer L: Test-retest reliability of the king-devick test in an adolescent population. J Athl Train. 2017; 52 (5): 439–45. PubMed Abstract | Publisher Full Text | Free Full Text 10. Goss DA, Becker E: Comparison of near fusional vergence ranges with rotary prisms and with prism bars. Optometry. 2011; 82 (2): 104–7. PubMed Abstract | Publisher Full Text 11. Wang J, Hatt SR, O'Connor AR, et al. : Final version of the Distance Randot Stereotest: normative data, reliability, and validity. J AAPOS. 2010; 14 (2): 142–6. PubMed Abstract | Publisher Full Text | Free Full Text 12. Pearce KL, Sufrinko A, Lau BC, et al. : Near Point of Convergence After a Sport-Related Concussion: Measurement Reliability and Relationship to Neurocognitive Impairment and Symptoms. Am J Sports Med. 2015; 43 (12): 3055–61. PubMed Abstract | Publisher Full Text | Free Full Text 13. Maples WC, Hoenes R: Near point of convergence norms measured in elementary school children. Optom Vis Sci. 2007; 84 (3): 224–8. PubMed Abstract | Publisher Full Text 14. Antona B, Barrio A, Barra F, et al. : Repeatability and agreement in the measurement of horizontal fusional vergences. Ophthalmic Physiol Opt. 2008; 28 (5): 475–91. PubMed Abstract | Publisher Full Text 15. Brozek J, Simonson E, Bushard W, et al. : Effects of practice and the consistency of repeated measurements of accommodation and vergence. Am J Ophthalmol. 1948; 31 (2): 191–8. PubMed Abstract | Publisher Full Text 16. Rouse MW, Borsting E, Deland PN, et al. : Reliability of binocular vision measurements used in the classification of convergence insufficiency. Optom Vis Sci. 2002; 79 (4): 254–64. PubMed Abstract | Publisher Full Text 17. Long S, Leclerc S, Tinjust D, et al. : Determining consistency and agreement of scores across two measurements of the visual system: Test-retest reliability. Med Sci Sports Exerc. 2018; 50 (5S): 664. Publisher Full Text 18. Giffard P, Daly L, Treleaven J: Influence of neck torsion on near point convergence in subjects with idiopathic neck pain. Musculoskelet Sci Pract. 2017; 32 : 51–6. PubMed Abstract | Publisher Full Text 19. Moganeswari D, Thomas J, Srinivasan K, et al. : Test Re-Test Reliability and Validity of Different Visual Acuity and Stereoacuity Charts Used in Preschool Children. J Clin Diagn Res. 2015; 9 (11): NC01–5. PubMed Abstract | Publisher Full Text | Free Full Text 20. Ettinger U, Kumari V, Crawford TJ, et al. : Reliability of smooth pursuit, fixation, and saccadic eye movements. Psychophysiology. 2003; 40 (4): 620–8. PubMed Abstract | Publisher Full Text 21. Klein C, Fischer B: Instrumental and test-retest reliability of saccadic measures. Biol Psychol. 2005; 68 (3): 201–213. PubMed Abstract | Publisher Full Text 22. Broglio SP, Ferrara MS, Macciocchi SN, et al. : Test-retest reliability of computerized concussion assessment programs. J Athl Train. 2007; 42 (4): 509–14. PubMed Abstract | Free Full Text 23. Pediatric Eye Disease Investigator Group: Interobserver reliability of the prism and alternate cover test in children with esotropia. Arch Ophthalmol. 2009; 127 (1): 59–65. PubMed Abstract | Publisher Full Text | Free Full Text 24. Stereo optical co: Randot stereotest. In: Stereo optical co., ed.; 1995. Reference Source 25. Rowe F: Clinical orthoptics. 3rd ed. Chichester, West Sussex: Wiley-Blackwell; 2012. Publisher Full Text 26. D'Agostino D: Basic examination: Physiology of eye movements - measurement of ductions, versions, and vergences. In: Scott W, D'Agostino D, Weingeist Lennarson L, editors. Orthoptics and ocular examination techniques . Baltimore: Williams & Wilkins; 1983. Reference Source 27. Hurtt J, Rasicovici A, Windsor C: Comprehensive review of orthoptics and ocular motility: Theory, therapy, and surgery. 2nd ed. Saint Louis: The C.V. Mosby Company; 1977. Reference Source 28. Bishop A: Convergence and convergent fusional reserves - investigation and treatment. In: Doshi S, Evans BJW, editors. Binocular vision and orthoptics: Investigation and management . Oxford: Butterworth-Heineman; 2001; 28–33. Publisher Full Text 29. Scheiman M, Gwiazda J, Li T: Non-surgical interventions for convergence insufficiency. Cochrane Database Syst Rev. 2011; (3): CD006768. PubMed Abstract | Publisher Full Text | Free Full Text 30. Hayes GJ, Cohen BE, Rouse MW, et al. : Normative values for the nearpoint of convergence of elementary schoolchildren. Optom Vis Sci. 1998; 75 (7): 506–12. PubMed Abstract | Publisher Full Text 31. Sutter P, Harvey L: Vision rehabilitation: Multidisciplinary care of the patient following brain injury. Boca Raton: Taylor & Francis Group; 2011. Reference Source 32. Birch E, Williams C, Drover J, et al. : Randot preschool stereoacuity test: Normative data and validity. J AAPOS. 2008; 12 (1): 23–6. PubMed Abstract | Publisher Full Text | Free Full Text 33. Piano ME, Tidbury LP, O'Connor AR: Normative Values for Near and Distance Clinical Tests of Stereoacuity. Strabismus. 2016; 24 (4): 169–72. PubMed Abstract | Publisher Full Text 34. Shrout PE, Fleiss JL: Intraclass correlations: uses in assessing rater reliability. Psychol Bull. 1979; 86 (2): 420–8. PubMed Abstract | Publisher Full Text 35. Bland JM, Altman D: Statistical methods for assessing agreement between two methods of clinical measurement. Lancet. 1986; 1 (8476): 307–10. PubMed Abstract | Publisher Full Text 36. Koo TK, Li MY: A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016; 15 (2): 155–63. PubMed Abstract | Publisher Full Text | Free Full Text 37. Johnston BC, Thorlund K, Schünemann HJ, et al. : Improving the interpretation of quality of life evidence in meta-analyses: the application of minimal important difference units. Health Qual Life Outcomes. 2010; 8 (1): 116. PubMed Abstract | Publisher Full Text | Free Full Text 38. Saville DJ: Multiple comparison procedures: The practical solution. Am Stat. 1990; 44 (2): 174–80. Publisher Full Text 39. R core team: R: A language and environment for statistical computing. Vienna, Austria: R foundation for statistical computing; 2015. 2015. Reference Source 40. Shrier I: Vision Tests in Concussion. 2019. http://www.doi.org/10.17605/OSF.IO/VB4W8 41. Moser RS, Schatz P, Grosner E, et al. : One year test-retest reliability of neurocognitive baseline scores in 10- to 12-year olds. Appl Neuropsychol Child. 2017; 6 (2): 166–71. PubMed Abstract | Publisher Full Text 42. de Jongh E, Leach C, Tjon-Fo-Sang M, et al. : Inter-examiner variability and agreement of the alternate prism cover test (APCT) measurements of strabismus performed by 4 examiners. Strabismus. 2014; 22 (4): 158–66. PubMed Abstract | Publisher Full Text 43. Mainwaring L, Pennock KMF, Mylabathula S, et al. : Subconcussive head impacts in sport: A systematic review of the evidence. Int J Psychophysiol. 2018; 132 (Pt A): 39–54. PubMed Abstract | Publisher Full Text 44. Bujang MA, Baharum N: A simplified guide to determination of sample size requirements for estimating the value of intraclass correlation coefficient: a review. Arch Orofac Sci. 2017; 12 (1): 1–11. Reference Source Comments on this article Comments (0) Version 5 VERSION 5 PUBLISHED 09 Jul 2019 ADD YOUR COMMENT Comment Author details Author details 1 Department of Epidemiology, Biostatistics and Occupational Health, McGill University, Montreal, Canada 2 Department of Health Research Methods, Evidence, and Impact, Michael G. DeGroote School of Medicine, McMaster University, Hamilton, Canada 3 Institut National du Sport du Quebec, Montreal, Canada 4 Department of Family Medicine, McGill University, Montreal, Canada 5 Faculty of Medicine, McGill University, Montreal, Canada 6 Department of Mathematics and Statistics, McGill University, Montreal, Canada 7 Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada Mehdi Aloosh Roles: Writing – Original Draft Preparation, Writing – Review & Editing Suzanne Leclerc Roles: Conceptualization, Funding Acquisition, Resources, Writing – Review & Editing Stephanie Long Roles: Data Curation, Writing – Review & Editing Guowei Zhong Roles: Data Curation, Writing – Review & Editing James M. Brophy Roles: Validation, Writing – Review & Editing Tibor Schuster Roles: Formal Analysis Russell Steele Roles: Formal Analysis, Writing – Review & Editing Ian Shrier Roles: Conceptualization, Formal Analysis, Funding Acquisition, Investigation, Resources, Supervision, Writing – Review & Editing Competing interests No competing interests were disclosed. Grant information This project was funded by government sources (MITACS [IT08159] and MEDTEQ [G245120], a non-profit organization (Institut National du Sport du Quebec) and private industry (APEXk Inc, and Varitron Inc.) The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Article Versions (5) version 5 Revised Published: 09 Sep 2020, 8:1032 https://doi.org/10.12688/f1000research.19587.5 version 4 Revised Published: 26 Aug 2020, 8:1032 https://doi.org/10.12688/f1000research.19587.4 version 3 Revised Published: 08 Jun 2020, 8:1032 https://doi.org/10.12688/f1000research.19587.3 version 2 Revised Published: 31 Mar 2020, 8:1032 https://doi.org/10.12688/f1000research.19587.2 version 1 Published: 09 Jul 2019, 8:1032 https://doi.org/10.12688/f1000research.19587.1 Copyright © 2020 Aloosh M et al . This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. Download Export To Sciwheel Bibtex EndNote ProCite Ref. Manager (RIS) Sente metrics Views Downloads F1000Research - - PubMed Central info_outline Data from PMC are received and updated monthly. - - Citations open_in_new 0 open_in_new 0 open_in_new SEE MORE DETAILS CITE how to cite this article Aloosh M, Leclerc S, Long S et al. One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.12688/f1000research.19587.5 ) NOTE: If applicable, it is important to ensure the information in square brackets after the title is included in all citations of this article. COPY CITATION DETAILS track receive updates on this article Track an article to receive email alerts on any updates to this article. TRACK THIS ARTICLE Share Open Peer Review Current Reviewer Status: ? Key to Reviewer Statuses VIEW HIDE Approved The paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved Fundamental flaws in the paper seriously undermine the findings and conclusions Version 5 VERSION 5 PUBLISHED 09 Sep 2020 Revised Views 0 Cite How to cite this report: Dickey JP and Richards D. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.29392.r71042 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v5#referee-response-71042 NOTE: it is important to ensure the information in square brackets after the title is included in this citation. Close Copy Citation Details Reviewer Report 10 Sep 2020 James P Dickey , School of Kinesiology, University of Western Ontario, London, Ontario, Canada Dillon Richards , Health and Rehabilitation Sciences, University of Western Ontario, London, Canada Approved VIEWS 0 https://doi.org/10.5256/f1000research.29392.r71042 The authors have revised the paper to acknowledge "Some data in these athletes appear to be outside the normative range of data previously described for the general population", and have provided access to the raw data through the data availability ... Continue reading READ ALL The authors have revised the paper to acknowledge "Some data in these athletes appear to be outside the normative range of data previously described for the general population", and have provided access to the raw data through the data availability link. This enables the readers to evaluate the credibility fo the data and interpret the findings accordingly. Competing Interests: No competing interests were disclosed. Reviewer Expertise: Biomechanics, head impact exposure in sports, concussion We confirm that we have read this submission and believe that we have an appropriate level of expertise to confirm that it is of an acceptable scientific standard. Close READ LESS CITE CITE HOW TO CITE THIS REPORT Dickey JP and Richards D. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.29392.r71042 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v5#referee-response-71042 NOTE: it is important to ensure the information in square brackets after the title is included in all citations of this article. COPY CITATION DETAILS Report a concern Respond or Comment COMMENT ON THIS REPORT Version 4 VERSION 4 PUBLISHED 26 Aug 2020 Revised Views 0 Cite How to cite this report: Richards D and Dickey JP. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.28661.r70272 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v4#referee-response-70272 NOTE: it is important to ensure the information in square brackets after the title is included in this citation. Close Copy Citation Details Reviewer Report 02 Sep 2020 Dillon Richards , Health and Rehabilitation Sciences, University of Western Ontario, London, Canada James P Dickey , School of Kinesiology, University of Western Ontario, London, Ontario, Canada Approved with Reservations VIEWS 0 https://doi.org/10.5256/f1000research.28661.r70272 Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points ... Continue reading READ ALL Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points are extreme outliers (four data points at 3 or more IQRs above the third quartile, including one value 8.125 IQRs above the third quartile), and that they exceed the range of normative data for Positive Fusional Vergence that you present in Table 1. Our previous review stated that there is a strong reason for believing that these data points are questionable. In fact there is strong evidence that these data points should be eliminated rather than simply evaluating their influence in a "sensitivity analysis". Competing Interests: No competing interests were disclosed. Reviewer Expertise: Biomechanics, head impact exposure in sports, concussion. We confirm that we have read this submission and believe that we have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however we have significant reservations, as outlined above. Close READ LESS CITE CITE HOW TO CITE THIS REPORT Richards D and Dickey JP. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.28661.r70272 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v4#referee-response-70272 NOTE: it is important to ensure the information in square brackets after the title is included in all citations of this article. COPY CITATION DETAILS Report a concern Author Response 09 Sep 2020 Ian Shrier , Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada 09 Sep 2020 Author Response Responses to reviewer 2&3: Comment: Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. ... Continue reading Responses to reviewer 2&3: Comment: Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points are extreme outliers (four data points at 3 or more IQRs above the third quartile, including one value 8.125 IQRs above the third quartile), and that they exceed the range of normative data for Positive Fusional Vergence that you present in Table 1. Our previous review stated that there is a strong reason for believing that these data points are questionable. In fact there is strong evidence that these data points should be eliminated rather than simply evaluating their influence in a “sensitivity analysis”. Response: We thank the reviewers for their feedback. There are three points raised. The reviewers insist that results excluding outliers be considered the primary analysis, and the results including all the data be considered secondary. The reviewers suggest there are four outliers instead of the 2 outliers we noted. The reviewers suggest that 4 points exceed the range of normative data we provided. 1. Sensitivity Analyses We think that there are different ways to look at outliers and interpret the results. As we mentioned in our previous response to the reviewers, we think that “eliminating” data, as the reviewers suggest, is not the optimal approach. Instead, the sensitivity analysis, as we performed, is a preferable approach unless there is a clear data error. In this way, we are transparent about our data set and our analysis, and the readers can evaluate our findings. F1000Research does not have an editor as an arbitrator when authors and reviewers disagree. Therefore, we have made the changes recommended by the reviewer and the analysis with the excluded data is now considered the primary result. We had already based our conclusions on the analyses after exclusions and have now edited the rest of the text as well. 2. Outliers We are not sure why the reviewer thinks there are four outliers. The data in our study are available online at https://osf.io/gnjdm/ . Here are the calculations for outliers, which we defined using the common standard: 1.5*IQR above the 3 rd quantile. Positive Fusional Vergence 30cm First Test: 25%: 20 75%: 31.25 IQR: 11.25 1.5*IQR: 16.9 Outlier Threshold (75%+1.5*IQR): 48.2 Second Test 25%: 23.75 75%: 30 IQR: 6.25 1.5*IQR: 9.4 Outlier Threshold (75%+1.5*IQR): 39.4 There is only one person with values that should be considered as outliers for positive fusional vergence at 30cm (Id=14). This occurred for both tests (90 on the first test and 85 on the second test). The text now reads: “Given the very high ICC and the presence of an outlier that greatly increased the range of the values for the measure (known to increase ICC), we repeated the analysis excluding the outlier. This decreased the ICC from 0.93 to 0.53, and increased the LoA to ±43.5%. One of the reviewers for this paper has insisted that the analysis without the outlier be considered the primary analysis.” We also added a sentence to the figure legend: “When the analysis was repeated excluding the outlier to the far right, the ICC decreased to 0.53 and the LoA increased to 43.5%.” Positive Fusional Vergence 3m First Test 25%: 17.5 75%: 25 IQR: 7.5 1.5*IQR: 11.3 Outlier Threshold (75%+1.5*IQR): 36.3 Second Test 25%: 17.5 75%: 21.25 IQR: 3.75 1.5*IQR: 5.6 Outlier Threshold (75%+1.5*IQR): 26.8 There are two people with values that should be considered as outliers for positive fusional vergence at 3m (Ids 14 and 15). Particiant 14 is an outlier for both measures, and Particicpant 15 is an outlier for the first test. When Participant 14 was removed the ICC dropped from 0.56 to 0.21 as we reported. If we remove only Participant 15, the ICC actually increases from 0.56 to 0.63. If we remove both outliers, the ICC was 0.45 and the LoA decreased from 60.2% to 41.4%. As per the reviewer’s request, we are reporting the analysis with both outliers removed. The text now reads: “There were two outliers for Positive Fusional Vergence at 3m (one participant on both measures and one participant on only one measure). When we removed both of these outliers, the ICC dropped from 0.56 to 0.45 and the LoA decreased from 60.2% to 41.4%. In both of these cases, the two scores from the outlier were quite different. Although one might anticipate that the ICC would increase by removing such outliers, the ICC actually decreased because the range of values for the measure decreased substantially. As above, one of the reviewers for this paper insisted that the analysis without the outliers be considered the primary analysis.” We have also added a sentence to the figure legend that says: “When the analysis for Positive Fusional Vergence at 3m was repeated excluding the two outliers, the ICC decreased to 0.45 and the LoA decreased to 41.4%.” The results of these tests are also reported in the Discussion. The text in that section now reads: “When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is similar to the value found for the one-week test-retest reliability (ICC=0.54); the LoA increased to 43.5%. When we removed the two outliers from Positive Fusional Vergence at 3m, the ICC decreased to 0.45 and LoA decreased to 41.4%. Note that the outliers for this measure had large differences between the two test scores, and removing such data points would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” 4. Normative Data Range We do not know why some data were outside previously described normative data range. These are the data we received from the clinician doing the test as part of his regular clinical practice. It is possible that previously published normative data for the population does not represent normative data for athletes like those included in our study. We have added one sentence mentioning this in the limitations section of the article. It says: “Some of the data in these athletes appear to be outside the normative range of data previously described for the general population.” Responses to reviewer 2&3: Comment: Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points are extreme outliers (four data points at 3 or more IQRs above the third quartile, including one value 8.125 IQRs above the third quartile), and that they exceed the range of normative data for Positive Fusional Vergence that you present in Table 1. Our previous review stated that there is a strong reason for believing that these data points are questionable. In fact there is strong evidence that these data points should be eliminated rather than simply evaluating their influence in a “sensitivity analysis”. Response: We thank the reviewers for their feedback. There are three points raised. The reviewers insist that results excluding outliers be considered the primary analysis, and the results including all the data be considered secondary. The reviewers suggest there are four outliers instead of the 2 outliers we noted. The reviewers suggest that 4 points exceed the range of normative data we provided. 1. Sensitivity Analyses We think that there are different ways to look at outliers and interpret the results. As we mentioned in our previous response to the reviewers, we think that “eliminating” data, as the reviewers suggest, is not the optimal approach. Instead, the sensitivity analysis, as we performed, is a preferable approach unless there is a clear data error. In this way, we are transparent about our data set and our analysis, and the readers can evaluate our findings. F1000Research does not have an editor as an arbitrator when authors and reviewers disagree. Therefore, we have made the changes recommended by the reviewer and the analysis with the excluded data is now considered the primary result. We had already based our conclusions on the analyses after exclusions and have now edited the rest of the text as well. 2. Outliers We are not sure why the reviewer thinks there are four outliers. The data in our study are available online at https://osf.io/gnjdm/ . Here are the calculations for outliers, which we defined using the common standard: 1.5*IQR above the 3 rd quantile. Positive Fusional Vergence 30cm First Test: 25%: 20 75%: 31.25 IQR: 11.25 1.5*IQR: 16.9 Outlier Threshold (75%+1.5*IQR): 48.2 Second Test 25%: 23.75 75%: 30 IQR: 6.25 1.5*IQR: 9.4 Outlier Threshold (75%+1.5*IQR): 39.4 There is only one person with values that should be considered as outliers for positive fusional vergence at 30cm (Id=14). This occurred for both tests (90 on the first test and 85 on the second test). The text now reads: “Given the very high ICC and the presence of an outlier that greatly increased the range of the values for the measure (known to increase ICC), we repeated the analysis excluding the outlier. This decreased the ICC from 0.93 to 0.53, and increased the LoA to ±43.5%. One of the reviewers for this paper has insisted that the analysis without the outlier be considered the primary analysis.” We also added a sentence to the figure legend: “When the analysis was repeated excluding the outlier to the far right, the ICC decreased to 0.53 and the LoA increased to 43.5%.” Positive Fusional Vergence 3m First Test 25%: 17.5 75%: 25 IQR: 7.5 1.5*IQR: 11.3 Outlier Threshold (75%+1.5*IQR): 36.3 Second Test 25%: 17.5 75%: 21.25 IQR: 3.75 1.5*IQR: 5.6 Outlier Threshold (75%+1.5*IQR): 26.8 There are two people with values that should be considered as outliers for positive fusional vergence at 3m (Ids 14 and 15). Particiant 14 is an outlier for both measures, and Particicpant 15 is an outlier for the first test. When Participant 14 was removed the ICC dropped from 0.56 to 0.21 as we reported. If we remove only Participant 15, the ICC actually increases from 0.56 to 0.63. If we remove both outliers, the ICC was 0.45 and the LoA decreased from 60.2% to 41.4%. As per the reviewer’s request, we are reporting the analysis with both outliers removed. The text now reads: “There were two outliers for Positive Fusional Vergence at 3m (one participant on both measures and one participant on only one measure). When we removed both of these outliers, the ICC dropped from 0.56 to 0.45 and the LoA decreased from 60.2% to 41.4%. In both of these cases, the two scores from the outlier were quite different. Although one might anticipate that the ICC would increase by removing such outliers, the ICC actually decreased because the range of values for the measure decreased substantially. As above, one of the reviewers for this paper insisted that the analysis without the outliers be considered the primary analysis.” We have also added a sentence to the figure legend that says: “When the analysis for Positive Fusional Vergence at 3m was repeated excluding the two outliers, the ICC decreased to 0.45 and the LoA decreased to 41.4%.” The results of these tests are also reported in the Discussion. The text in that section now reads: “When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is similar to the value found for the one-week test-retest reliability (ICC=0.54); the LoA increased to 43.5%. When we removed the two outliers from Positive Fusional Vergence at 3m, the ICC decreased to 0.45 and LoA decreased to 41.4%. Note that the outliers for this measure had large differences between the two test scores, and removing such data points would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” 4. Normative Data Range We do not know why some data were outside previously described normative data range. These are the data we received from the clinician doing the test as part of his regular clinical practice. It is possible that previously published normative data for the population does not represent normative data for athletes like those included in our study. We have added one sentence mentioning this in the limitations section of the article. It says: “Some of the data in these athletes appear to be outside the normative range of data previously described for the general population.” Competing Interests: No competing interests were disclosed. Close Report a concern Respond or Comment COMMENTS ON THIS REPORT Author Response 09 Sep 2020 Ian Shrier , Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada 09 Sep 2020 Author Response Responses to reviewer 2&3: Comment: Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. ... Continue reading Responses to reviewer 2&3: Comment: Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points are extreme outliers (four data points at 3 or more IQRs above the third quartile, including one value 8.125 IQRs above the third quartile), and that they exceed the range of normative data for Positive Fusional Vergence that you present in Table 1. Our previous review stated that there is a strong reason for believing that these data points are questionable. In fact there is strong evidence that these data points should be eliminated rather than simply evaluating their influence in a “sensitivity analysis”. Response: We thank the reviewers for their feedback. There are three points raised. The reviewers insist that results excluding outliers be considered the primary analysis, and the results including all the data be considered secondary. The reviewers suggest there are four outliers instead of the 2 outliers we noted. The reviewers suggest that 4 points exceed the range of normative data we provided. 1. Sensitivity Analyses We think that there are different ways to look at outliers and interpret the results. As we mentioned in our previous response to the reviewers, we think that “eliminating” data, as the reviewers suggest, is not the optimal approach. Instead, the sensitivity analysis, as we performed, is a preferable approach unless there is a clear data error. In this way, we are transparent about our data set and our analysis, and the readers can evaluate our findings. F1000Research does not have an editor as an arbitrator when authors and reviewers disagree. Therefore, we have made the changes recommended by the reviewer and the analysis with the excluded data is now considered the primary result. We had already based our conclusions on the analyses after exclusions and have now edited the rest of the text as well. 2. Outliers We are not sure why the reviewer thinks there are four outliers. The data in our study are available online at https://osf.io/gnjdm/ . Here are the calculations for outliers, which we defined using the common standard: 1.5*IQR above the 3 rd quantile. Positive Fusional Vergence 30cm First Test: 25%: 20 75%: 31.25 IQR: 11.25 1.5*IQR: 16.9 Outlier Threshold (75%+1.5*IQR): 48.2 Second Test 25%: 23.75 75%: 30 IQR: 6.25 1.5*IQR: 9.4 Outlier Threshold (75%+1.5*IQR): 39.4 There is only one person with values that should be considered as outliers for positive fusional vergence at 30cm (Id=14). This occurred for both tests (90 on the first test and 85 on the second test). The text now reads: “Given the very high ICC and the presence of an outlier that greatly increased the range of the values for the measure (known to increase ICC), we repeated the analysis excluding the outlier. This decreased the ICC from 0.93 to 0.53, and increased the LoA to ±43.5%. One of the reviewers for this paper has insisted that the analysis without the outlier be considered the primary analysis.” We also added a sentence to the figure legend: “When the analysis was repeated excluding the outlier to the far right, the ICC decreased to 0.53 and the LoA increased to 43.5%.” Positive Fusional Vergence 3m First Test 25%: 17.5 75%: 25 IQR: 7.5 1.5*IQR: 11.3 Outlier Threshold (75%+1.5*IQR): 36.3 Second Test 25%: 17.5 75%: 21.25 IQR: 3.75 1.5*IQR: 5.6 Outlier Threshold (75%+1.5*IQR): 26.8 There are two people with values that should be considered as outliers for positive fusional vergence at 3m (Ids 14 and 15). Particiant 14 is an outlier for both measures, and Particicpant 15 is an outlier for the first test. When Participant 14 was removed the ICC dropped from 0.56 to 0.21 as we reported. If we remove only Participant 15, the ICC actually increases from 0.56 to 0.63. If we remove both outliers, the ICC was 0.45 and the LoA decreased from 60.2% to 41.4%. As per the reviewer’s request, we are reporting the analysis with both outliers removed. The text now reads: “There were two outliers for Positive Fusional Vergence at 3m (one participant on both measures and one participant on only one measure). When we removed both of these outliers, the ICC dropped from 0.56 to 0.45 and the LoA decreased from 60.2% to 41.4%. In both of these cases, the two scores from the outlier were quite different. Although one might anticipate that the ICC would increase by removing such outliers, the ICC actually decreased because the range of values for the measure decreased substantially. As above, one of the reviewers for this paper insisted that the analysis without the outliers be considered the primary analysis.” We have also added a sentence to the figure legend that says: “When the analysis for Positive Fusional Vergence at 3m was repeated excluding the two outliers, the ICC decreased to 0.45 and the LoA decreased to 41.4%.” The results of these tests are also reported in the Discussion. The text in that section now reads: “When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is similar to the value found for the one-week test-retest reliability (ICC=0.54); the LoA increased to 43.5%. When we removed the two outliers from Positive Fusional Vergence at 3m, the ICC decreased to 0.45 and LoA decreased to 41.4%. Note that the outliers for this measure had large differences between the two test scores, and removing such data points would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” 4. Normative Data Range We do not know why some data were outside previously described normative data range. These are the data we received from the clinician doing the test as part of his regular clinical practice. It is possible that previously published normative data for the population does not represent normative data for athletes like those included in our study. We have added one sentence mentioning this in the limitations section of the article. It says: “Some of the data in these athletes appear to be outside the normative range of data previously described for the general population.” Responses to reviewer 2&3: Comment: Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points are extreme outliers (four data points at 3 or more IQRs above the third quartile, including one value 8.125 IQRs above the third quartile), and that they exceed the range of normative data for Positive Fusional Vergence that you present in Table 1. Our previous review stated that there is a strong reason for believing that these data points are questionable. In fact there is strong evidence that these data points should be eliminated rather than simply evaluating their influence in a “sensitivity analysis”. Response: We thank the reviewers for their feedback. There are three points raised. The reviewers insist that results excluding outliers be considered the primary analysis, and the results including all the data be considered secondary. The reviewers suggest there are four outliers instead of the 2 outliers we noted. The reviewers suggest that 4 points exceed the range of normative data we provided. 1. Sensitivity Analyses We think that there are different ways to look at outliers and interpret the results. As we mentioned in our previous response to the reviewers, we think that “eliminating” data, as the reviewers suggest, is not the optimal approach. Instead, the sensitivity analysis, as we performed, is a preferable approach unless there is a clear data error. In this way, we are transparent about our data set and our analysis, and the readers can evaluate our findings. F1000Research does not have an editor as an arbitrator when authors and reviewers disagree. Therefore, we have made the changes recommended by the reviewer and the analysis with the excluded data is now considered the primary result. We had already based our conclusions on the analyses after exclusions and have now edited the rest of the text as well. 2. Outliers We are not sure why the reviewer thinks there are four outliers. The data in our study are available online at https://osf.io/gnjdm/ . Here are the calculations for outliers, which we defined using the common standard: 1.5*IQR above the 3 rd quantile. Positive Fusional Vergence 30cm First Test: 25%: 20 75%: 31.25 IQR: 11.25 1.5*IQR: 16.9 Outlier Threshold (75%+1.5*IQR): 48.2 Second Test 25%: 23.75 75%: 30 IQR: 6.25 1.5*IQR: 9.4 Outlier Threshold (75%+1.5*IQR): 39.4 There is only one person with values that should be considered as outliers for positive fusional vergence at 30cm (Id=14). This occurred for both tests (90 on the first test and 85 on the second test). The text now reads: “Given the very high ICC and the presence of an outlier that greatly increased the range of the values for the measure (known to increase ICC), we repeated the analysis excluding the outlier. This decreased the ICC from 0.93 to 0.53, and increased the LoA to ±43.5%. One of the reviewers for this paper has insisted that the analysis without the outlier be considered the primary analysis.” We also added a sentence to the figure legend: “When the analysis was repeated excluding the outlier to the far right, the ICC decreased to 0.53 and the LoA increased to 43.5%.” Positive Fusional Vergence 3m First Test 25%: 17.5 75%: 25 IQR: 7.5 1.5*IQR: 11.3 Outlier Threshold (75%+1.5*IQR): 36.3 Second Test 25%: 17.5 75%: 21.25 IQR: 3.75 1.5*IQR: 5.6 Outlier Threshold (75%+1.5*IQR): 26.8 There are two people with values that should be considered as outliers for positive fusional vergence at 3m (Ids 14 and 15). Particiant 14 is an outlier for both measures, and Particicpant 15 is an outlier for the first test. When Participant 14 was removed the ICC dropped from 0.56 to 0.21 as we reported. If we remove only Participant 15, the ICC actually increases from 0.56 to 0.63. If we remove both outliers, the ICC was 0.45 and the LoA decreased from 60.2% to 41.4%. As per the reviewer’s request, we are reporting the analysis with both outliers removed. The text now reads: “There were two outliers for Positive Fusional Vergence at 3m (one participant on both measures and one participant on only one measure). When we removed both of these outliers, the ICC dropped from 0.56 to 0.45 and the LoA decreased from 60.2% to 41.4%. In both of these cases, the two scores from the outlier were quite different. Although one might anticipate that the ICC would increase by removing such outliers, the ICC actually decreased because the range of values for the measure decreased substantially. As above, one of the reviewers for this paper insisted that the analysis without the outliers be considered the primary analysis.” We have also added a sentence to the figure legend that says: “When the analysis for Positive Fusional Vergence at 3m was repeated excluding the two outliers, the ICC decreased to 0.45 and the LoA decreased to 41.4%.” The results of these tests are also reported in the Discussion. The text in that section now reads: “When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is similar to the value found for the one-week test-retest reliability (ICC=0.54); the LoA increased to 43.5%. When we removed the two outliers from Positive Fusional Vergence at 3m, the ICC decreased to 0.45 and LoA decreased to 41.4%. Note that the outliers for this measure had large differences between the two test scores, and removing such data points would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” 4. Normative Data Range We do not know why some data were outside previously described normative data range. These are the data we received from the clinician doing the test as part of his regular clinical practice. It is possible that previously published normative data for the population does not represent normative data for athletes like those included in our study. We have added one sentence mentioning this in the limitations section of the article. It says: “Some of the data in these athletes appear to be outside the normative range of data previously described for the general population.” Competing Interests: No competing interests were disclosed. Close Report a concern COMMENT ON THIS REPORT Views 0 Cite How to cite this report: Haider MN. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.28661.r70273 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v4#referee-response-70273 NOTE: it is important to ensure the information in square brackets after the title is included in this citation. Close Copy Citation Details Reviewer Report 27 Aug 2020 M Nadir Haider , Jacobs School of Medicine and Biomedical Sciences, State University of New York at Buffalo, Buffalo, NY, USA Approved VIEWS 0 https://doi.org/10.5256/f1000research.28661.r70273 No ... Continue reading READ ALL No additional comments. Competing Interests: No competing interests were disclosed. Reviewer Expertise: Concussion, biomarkers, physiology, cerebral blood flow. I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard. Close READ LESS CITE CITE HOW TO CITE THIS REPORT Haider MN. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.28661.r70273 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v4#referee-response-70273 NOTE: it is important to ensure the information in square brackets after the title is included in all citations of this article. COPY CITATION DETAILS Report a concern Respond or Comment COMMENT ON THIS REPORT Version 3 VERSION 3 PUBLISHED 08 Jun 2020 Revised Views 0 Cite How to cite this report: Richards D and Dickey JP. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.27076.r66529 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v3#referee-response-66529 NOTE: it is important to ensure the information in square brackets after the title is included in this citation. Close Copy Citation Details Reviewer Report 20 Jul 2020 Dillon Richards , Health and Rehabilitation Sciences, University of Western Ontario, London, Canada James P Dickey , School of Kinesiology, University of Western Ontario, London, Ontario, Canada Approved with Reservations VIEWS 0 https://doi.org/10.5256/f1000research.27076.r66529 This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. The outlying data points in Positive Fusional Vergence at 30 cm and 3 ... Continue reading READ ALL This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. The outlying data points in Positive Fusional Vergence at 30 cm and 3 m have been identified by the previous external peer reviewers and warrant additional consideration. You present your ICC findings with and without the outlying data points, which is appropriate. However, you do not fully characterize the extreme deviance of the outlying data points. Your wording about the outliers is rather misleading – you state “we noticed one outlier that greatly increased the range of values along x-axis in Figure 2 and Figure 3”. However, both the x- and y-coordinates of the outlier in Figure 2 meet your definition of outlier (1.5 interquartile ranges below the first quartile or above the third quartile), so it is not merely an issue with the x-axis. As well, you state “There was also one outlier for Positive Fusional Vergence at 3 m, 1.5 interquartile range above the third quartile”, but the 1.5 interquartile range (IQR) is your threshold for identifying outliers, not the description of the outlier – this data point is actually 8.125 IQRs above the third quartile. Statistically speaking, it is extremely unlikely that this data point is part of the same distribution as the rest of the data set. Finally, and perhaps most importantly, you state that you “have no reason to believe the data are inaccurate”. However, these outlying data points all exceed the range of normative data for Positive Fusional Vergence that you present in Table 1, providing a strong reason for believing that these data points are questionable. Your previous response “if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point” trivializes the issue. The issues about the outlier data points must be more thoroughly addressed in the manuscript. It is highly unfortunate that the sample size is so limited, particularly since it would appear that your inclusion criteria were quite broad (followed by the Institut National du Sport du Quebec from 2015–2018). Examination of the participants' durations between tests reveals that the majority of the participants had 335-336 or 371-371 days between assessments - presumably these dates correspond to the timing of the preseason tests for the different sports. Would you have more eligible participants if you had broadened the eligibility criterion? It is unclear how it could be that your participants were limited to waterpolo and short-track speed skating, when presumably you started with a larger number of sports, but this should be clarified as it may reflect a bias in participant selection. As well, both waterpolo (Black et al. 2017) 1 and short-track speed skating (Quinn et al. 2003) 2 have a relatively high rate of concussions, and presumably the athletes may have received subconcussive head impacts, without receiving a concussion. Repetitive hits to the head are associated with microstructural and functional changes in the brain (Mainwaring et al. 2018) 3 , and therefore should be acknowledged as a potential factor for the participants in this paper. You identify that test-retest reliability of vision tests has been evaluated at the 1 day to 45 days time span. However, studies have evaluated longer-term test-retest reliability. For example, Klein and Fischer (2005) 4 evaluated 19-month test–retest correlations of pro- and anti-saccadic eye movements on 117 participants. Of more direct relevance to the student athletes evaluated in your paper, Breedlove et al. (2019) 5 evaluated the reliability of the King-Devick test (prosaccades) on NCAA athletes, including 833 participants with measures one year apart, and Naidu et al. (2018) 6 evaluated the season-to-season reliability of the King-Devick Test in Canadian professional football players. Your paper would be strengthened by incorporating a fuller complement of relevant papers that have performed longer-term test-retest reliability measures of vision tests, and comparing your findings with theirs. The saccade measures reported in the paper have extremely limited value as they were collected using proprietary equipment - they should likely be removed from the paper. The scatterplots (Figures 2A, 3A and 4A) show the line of identity, but it would be interesting to also see the line of best fit. Furthermore, for the parameters with outliers, it would be interesting to add the lines of best fit with and without the outlier. The raw data presented through the Data Availability link is very helpful for gaining insight into the specifics of your data. However, it reveals that all of the data are reported as integers. Is this level of precision adequate for capturing the various vision tests? It would be helpful to include a "data dictionary", as recommended for best practices with spreadsheets (Broman and Woo, 2017) 7 . Is the work clearly and accurately presented and does it cite the current literature? Yes Is the study design appropriate and is the work technically sound? Partly Are sufficient details of methods and analysis provided to allow replication by others? Partly If applicable, is the statistical analysis and its interpretation appropriate? Yes Are all the source data underlying the results available to ensure full reproducibility? Yes Are the conclusions drawn adequately supported by the results? Partly References 1. Black AM, Sergio LE, Macpherson AK: The Epidemiology of Concussions: Number and Nature of Concussions and Time to Recovery Among Female and Male Canadian Varsity Athletes 2008 to 2011. Clin J Sport Med . 2017; 27 (1): 52-56 PubMed Abstract | Publisher Full Text 2. Quinn A, Lun V, McCall J, Overend T: Injuries in short track speed skating. Am J Sports Med . 31 (4): 507-10 PubMed Abstract | Publisher Full Text 3. Mainwaring L, Ferdinand Pennock KM, Mylabathula S, Alavie BZ: Subconcussive head impacts in sport: A systematic review of the evidence. Int J Psychophysiol . 132 (Pt A): 39-54 PubMed Abstract | Publisher Full Text 4. Klein C, Fischer B: Instrumental and test-retest reliability of saccadic measures. Biol Psychol . 2005; 68 (3): 201-13 PubMed Abstract | Publisher Full Text 5. Breedlove KM, Ortega JD, Kaminski TW, Harmon KG, et al.: King-Devick Test Reliability in National Collegiate Athletic Association Athletes: A National Collegiate Athletic Association-Department of Defense Concussion Assessment, Research and Education Report. J Athl Train . 2019; 54 (12): 1241-1246 PubMed Abstract | Publisher Full Text 6. Naidu D, Borza C, Kobitowich T, Mrazik M: Sideline Concussion Assessment: The King-Devick Test in Canadian Professional Football. Journal of Neurotrauma . 2018; 35 (19): 2283-2286 Publisher Full Text 7. Broman K, Woo K: Data Organization in Spreadsheets. The American Statistician . 2018; 72 (1): 2-10 Publisher Full Text Competing Interests: No competing interests were disclosed. Reviewer Expertise: Biomechanics, head impact exposure in sports, concussion. We confirm that we have read this submission and believe that we have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however we have significant reservations, as outlined above. Close READ LESS CITE CITE HOW TO CITE THIS REPORT Richards D and Dickey JP. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.27076.r66529 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v3#referee-response-66529 NOTE: it is important to ensure the information in square brackets after the title is included in all citations of this article. COPY CITATION DETAILS Report a concern Author Response 26 Aug 2020 Ian Shrier , Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada 26 Aug 2020 Author Response Author Responses REVIEWER # 2&3 Comment: This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. Answer: We thank the ... Continue reading Author Responses REVIEWER # 2&3 Comment: This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. Answer: We thank the reviewers for their interest in our manuscript. __________________ Comment: The outlying data points in Positive Fusional Vergence at 30 cm and 3 m have been identified by the previous external peer reviewers and warrant additional consideration. You present your ICC findings with and without the outlying data points, which is appropriate. However, you do not fully characterize the extreme deviance of the outlying data points. Your wording about the outliers is rather misleading – you state “we noticed one outlier that greatly increased the range of values along x-axis in Figure 2 and Figure 3”. However, both the x- and y-coordinates of the outlier in Figure 2 meet your definition of outlier (1.5 interquartile ranges below the first quartile or above the third quartile), so it is not merely an issue with the x-axis. As well, you state “There was also one outlier for Positive Fusional Vergence at 3 m, 1.5 interquartile range above the third quartile”, but the 1.5 interquartile range (IQR) is your threshold for identifying outliers, not the description of the outlier – this data point is actually 8.125 IQRs above the third quartile. Statistically speaking, it is extremely unlikely that this data point is part of the same distribution as the rest of the data set. Finally, and perhaps most importantly, you state that you “have no reason to believe the data are inaccurate”. However, these outlying data points all exceed the range of normative data for Positive Fusional Vergence that you present in Table 1, providing a strong reason for believing that these data points are questionable. Your previous response “if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point” trivializes the issue. The issues about the outlier data points must be more thoroughly addressed in the manuscript. Answer: We thank the reviewers for raising this point and have modified the text accordingly. For the comment that 1.5 IQR is the threshold and not the data point, we agree and have removed the phrase. It now reads: “There was also one outlier for Positive Fusional Vergence at 3m. When removing this outlier in a sensitivity analysis, the ICC dropped from 0.57 to 0.21.” The paragraph in which this is mentioned refers to the fact that the 1-year test-retest reliability had higher ICC than the 1-week test-retest reliability and this should not be possible. Our sensitivity analyses were conducted to determine if this occurred because of the increased range observed in the 1-year data. The reviewers are correct that the outlier in question is indeed an outlier on both the x and y axis. We have modified the text accordingly. This particular section now reads as below. Similar changes were made to other parts of the manuscript where appropriate: “Given the very high ICC and the presence of an outlier that greatly increased the range of values for the measure (known to increase ICC), we conducted a sensitivity analysis excluding the outlier.” With respect to justifying keeping the outlier in the plot or not, we did not mean to trivialize the issue. We only meant that the decision to remove an outlier needs more justification than simply that the data point was unexpected. Therefore, although we agree with the reviewers that our sensitivity analysis for the ICC is more likely to be correct, we do not feel there is enough evidence to replace the original analysis with the sensitivity analysis as the primary analysis. We feel that discussing this at length would be more confusing than helpful and have deleted the phrase related to “have no reason to believe the data are inaccurate”. The full paragraph now reads: “In one-year test-retest, Positive Fusional Vergence showed excellent reliability at 30cm (ICC=0.93) and moderate at 3m (ICC=0.56), initially. These values were better than the one-week test-retest reliability (ICC=0.54 and 0.49, respectively) 17 . It is difficult to understand how test-retest reliability over one year could be better than test-retest reliability over one week. When we explored the data further, we noticed one outlier that greatly increased the range of values for Positive Fusional Vergence at 30cm (Figure 2) and Positive Fusional Vergence at 3m (Figure 3). Increasing the range of values is known to increase the ICC. This is because ICC is based on the results of an analysis of variance which separates the error into variability between individuals (range of values along x or y axes) and variability within an individual. Therefore, if variability between persons increases, indicated by a larger range of values, ICC will increase. We explored how removing the outlier in our data would affect the results. When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is below the value found for the one-week test-retest reliability; it did not affect LoA. When we removed the outlier (same person) from Positive Fusional Vergence at 3m, the ICC decreased to 0.21. Note that the outlier for this measure had a large difference between the two test scores, and removing such a data point would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” __________________ Comment: It is highly unfortunate that the sample size is so limited, particularly since it would appear that your inclusion criteria were quite broad (followed by the Institut National du Sport du Quebec from 2015–2018). Examination of the participants' durations between tests reveals that the majority of the participants had 335-336 or 371-371 days between assessments - presumably these dates correspond to the timing of the preseason tests for the different sports. Would you have more eligible participants if you had broadened the eligibility criterion? Answer: We thank the reviewers for this comment. Our eligibility criteria only required that the athlete not have a concussion or undergo vision training between tests, and did not have a condition that would affect the results of vision testing. We are not sure which of these criteria the reviewers think we could relax and still obtain an unbiased answer to the question of 1-year test-retest reliability. We could have shortened the interval to only several months, but that would no longer be answering the 1-year test-retest reliability question. We have not made any changes to the manuscript. __________________ Comment: It is unclear how it could be that your participants were limited to waterpolo and short-track speed skating, when presumably you started with a larger number of sports, but this should be clarified as it may reflect a bias in participant selection. As well, both waterpolo (Black et al. 2017) and short-track speed skating (Quinn et al. 2003) have a relatively high rate of concussions, and presumably the athletes may have received subconcussive head impacts, without receiving a concussion. Repetitive hits to the head are associated with microstructural and functional changes in the brain (Mainwaring et al. 2018), and therefore should be acknowledged as a potential factor for the participants in this paper. Answer: We thank the reviewers for raising these points. For the types of sports participants were engaged in, these are the data provided to us. Many athletes from other sports only had 1 test, and some had concussions or vision testing within the 1-year interval. We do not have data on which athletes were referred for testing but never went for the test. The reviewers suggested Black et al reported waterpolo as a sport with many concussions. However, the study cited actually reported 0 concussions in waterpolo athletes. The Quinn et al study reported 6 concussions in 63 athletes over a 1-year period. The authors did not include the injury rate in the paper and it is not possible to compare risks to other sports without knowing how often the athletes were competing / practicing. In general, short-track speed skating concussions occur because of collisions that cause the athlete to fall, and then they may hit their head into the padded boards or on the ice. There are not multiple small hits like one would receive in American football or hockey. That said, we expand on the issue below for other studies that might include athletes from these types of sports. The reviewers suggest cumulative subconcussive head impacts should be raised as potential factor for participants in this study. We are not sure what the reviewers mean. We agree with the paper by Mainwaring et al. (2018) that the reviewers cited. Mainwaring et al states: “Both the research and conceptual understanding of this phenomenon are in their infancy” “the findings are equivocal regarding the effect of subconcussive impacts on the brain” “Insufficient evidence was presented to conclude that repetitive head impacts are associated with neurocognitive impairment. It may be that neuropsychological assessment tools are not sufficiently sensitive to detect any subtle changes in cognitive function that emerge from subconcussive impacts, or that the neurocognitive changes are inconsequential, or follow neurophysiological changes or damage.” “Future research is needed to characterize the phenomenon in question.” As an example, one study found that repetitive subconcussive head impacts over a single season do not appear to result in short-term neurologic impairment (see Gysland SM, Mihalik JP, Register-Mihalik JK, Trulock SC, Shields EW, Guskiewicz KM. The relationship between subconcussive impacts and concussion history on clinical measures of neurologic function in collegiate football players . Annals of biomedical engineering. 2012;40(1):14-22). Aside from these results that do not support a decrease in neurocognitive function with subconcussive impacts, our objective in this study was to report on the 1-year test-retest reliability of vision tests in order to help clinicians understand how to interpret differences between testing conducted post-concussion and at baseline. If subconcussive impacts did affect vision testing, one would expect a decline in visual function as a consequence of subconcussive impacts. If this occurred, any change in test scores between baseline and post-concussion could not be attributed to concussion. That said, we doubt this is the case because if that were true, one would expect a decline in vision function (and test scores) conducted one year after baseline. We did not observe this in our data. However, we acknowledge that the athletes in our study were not involved in sports with many subconcussive impacts. We have modified the text at the end of the limitation, which now includes: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm). In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. Despite starting with athletes from many sports, only athletes from Waterpolo and Short-track speed skating met our eligibility criteria. It is unclear if subconcussion impacts affect neurological function in general 43 . If subconcussion impacts were common in these sports and affected vision testing, we should have seen a systematic decrease in vision capacity between the two tests; this was not observed. Further, if it were present, the effect would be considered part of the “noise” clinicians have to consider when comparing the results from post-concussion and baseline tests.” __________________ Comment: You identify that test-retest reliability of vision tests has been evaluated at the 1-45 days time span. However, studies have evaluated longer-term test-retest reliability. For example, Klein and Fischer (2005) evaluated 19-month test–retest correlations of pro- and anti-saccadic eye movements on 117 participants. Of more direct relevance to the student athletes evaluated in your paper, Breedlove et al. (2019) evaluated the reliability of the King-Devick test (prosaccades) on NCAA athletes, including 833 participants with measures one year apart, and Naidu et al. (2018) evaluated the season-to-season reliability of the King-Devick Test in Canadian professional football players. Your paper would be strengthened by incorporating a fuller complement of relevant papers that have performed longer-term test-retest reliability measures of vision tests, and comparing your findings with theirs. Answer: We thank the reviewers for the reference that we had not been aware of. Our study investigated tests for specific visual function. Although the King-Devick test is sometimes used in concussion, it measures a combination of functions much beyond visual function. Therefore, we do not feel it is relevant to our research questions. We were not aware of the Klein and Fischer article and have now included the reference in the Introduction and Discussion. Our test was quite different from that studied in Klein and Fisher. The introduction text now reads: Previous investigations of the test-retest reliability of these vision tests have used short test-retest time intervals ranging from 1 day to 45 days 9 – 17 , except for one test of saccades 44 . and the Discussion text now reads: “However, we could not find any research examining the stability of the vision tests over a one year period, in athlete or non-athlete populations except for one test of saccades that was very different from the test used in this study 44 .” __________________ Comment: The saccade measures reported in the paper have extremely limited value as they were collected using proprietary equipment - they should likely be removed from the paper. Answer: We respectfully disagree with the reviewers. First, we do not see any harm in including the result of a non-standard test and readers who are not interested can simply ignore the results. Second, we evaluated this non-standard test as this measure was in our a priori protocol. Omitting analyses described in an a priori protocol is a form of reporting bias that we would prefer to avoid. We have modified the limitation section to say: “Finally, the results of the test of Saccades in this study are based on the unpublished proprietary algorithm developed by the clinician. This limits its applicability for other clinicians.” __________________ Comment: The scatterplots (Figures 2A, 3A and 4A) show the line of identity, but it would be interesting to also see the line of best fit. Furthermore, for the parameters with outliers, it would be interesting to add the lines of best fit with and without the outlier. Answer: We thank the reviewers for this comment. We believe the recommended statistical practice for evaluating reliability is the ICC with line of identity, and LOA. We have provided the references that guided this decision. Lines of best fit are not measures of reliability. In addition, any comparison of regression lines with the line of identity can be misleading because one must incorporate the uncertainty due to sampling. If the reviewers have an appropriate statistical reference that supports using regression in studies of test-retest reliability, we would be happy to add the analyses in a subsequent revision. __________________ Comment: The raw data presented through the Data Availability link is very helpful for gaining insight into the specifics of your data. However, it reveals that all of the data are reported as integers. Is this level of precision adequate for capturing the various vision tests? It would be helpful to include a "data dictionary", as recommended for best practices with spreadsheets (Broman and Woo, 2017). Answer: We thank the reviewers for their comment. We have developed a data dictionary and uploaded it as metadata. We agree that some of the measures could have been measured more precisely than others but these are the data provided from the clinician with expertise in orthoptics. We have added text to the beginning of the 2 nd paragraph in the limitations section which now reads: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm).” Author Responses REVIEWER # 2&3 Comment: This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. Answer: We thank the reviewers for their interest in our manuscript. __________________ Comment: The outlying data points in Positive Fusional Vergence at 30 cm and 3 m have been identified by the previous external peer reviewers and warrant additional consideration. You present your ICC findings with and without the outlying data points, which is appropriate. However, you do not fully characterize the extreme deviance of the outlying data points. Your wording about the outliers is rather misleading – you state “we noticed one outlier that greatly increased the range of values along x-axis in Figure 2 and Figure 3”. However, both the x- and y-coordinates of the outlier in Figure 2 meet your definition of outlier (1.5 interquartile ranges below the first quartile or above the third quartile), so it is not merely an issue with the x-axis. As well, you state “There was also one outlier for Positive Fusional Vergence at 3 m, 1.5 interquartile range above the third quartile”, but the 1.5 interquartile range (IQR) is your threshold for identifying outliers, not the description of the outlier – this data point is actually 8.125 IQRs above the third quartile. Statistically speaking, it is extremely unlikely that this data point is part of the same distribution as the rest of the data set. Finally, and perhaps most importantly, you state that you “have no reason to believe the data are inaccurate”. However, these outlying data points all exceed the range of normative data for Positive Fusional Vergence that you present in Table 1, providing a strong reason for believing that these data points are questionable. Your previous response “if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point” trivializes the issue. The issues about the outlier data points must be more thoroughly addressed in the manuscript. Answer: We thank the reviewers for raising this point and have modified the text accordingly. For the comment that 1.5 IQR is the threshold and not the data point, we agree and have removed the phrase. It now reads: “There was also one outlier for Positive Fusional Vergence at 3m. When removing this outlier in a sensitivity analysis, the ICC dropped from 0.57 to 0.21.” The paragraph in which this is mentioned refers to the fact that the 1-year test-retest reliability had higher ICC than the 1-week test-retest reliability and this should not be possible. Our sensitivity analyses were conducted to determine if this occurred because of the increased range observed in the 1-year data. The reviewers are correct that the outlier in question is indeed an outlier on both the x and y axis. We have modified the text accordingly. This particular section now reads as below. Similar changes were made to other parts of the manuscript where appropriate: “Given the very high ICC and the presence of an outlier that greatly increased the range of values for the measure (known to increase ICC), we conducted a sensitivity analysis excluding the outlier.” With respect to justifying keeping the outlier in the plot or not, we did not mean to trivialize the issue. We only meant that the decision to remove an outlier needs more justification than simply that the data point was unexpected. Therefore, although we agree with the reviewers that our sensitivity analysis for the ICC is more likely to be correct, we do not feel there is enough evidence to replace the original analysis with the sensitivity analysis as the primary analysis. We feel that discussing this at length would be more confusing than helpful and have deleted the phrase related to “have no reason to believe the data are inaccurate”. The full paragraph now reads: “In one-year test-retest, Positive Fusional Vergence showed excellent reliability at 30cm (ICC=0.93) and moderate at 3m (ICC=0.56), initially. These values were better than the one-week test-retest reliability (ICC=0.54 and 0.49, respectively) 17 . It is difficult to understand how test-retest reliability over one year could be better than test-retest reliability over one week. When we explored the data further, we noticed one outlier that greatly increased the range of values for Positive Fusional Vergence at 30cm (Figure 2) and Positive Fusional Vergence at 3m (Figure 3). Increasing the range of values is known to increase the ICC. This is because ICC is based on the results of an analysis of variance which separates the error into variability between individuals (range of values along x or y axes) and variability within an individual. Therefore, if variability between persons increases, indicated by a larger range of values, ICC will increase. We explored how removing the outlier in our data would affect the results. When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is below the value found for the one-week test-retest reliability; it did not affect LoA. When we removed the outlier (same person) from Positive Fusional Vergence at 3m, the ICC decreased to 0.21. Note that the outlier for this measure had a large difference between the two test scores, and removing such a data point would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” __________________ Comment: It is highly unfortunate that the sample size is so limited, particularly since it would appear that your inclusion criteria were quite broad (followed by the Institut National du Sport du Quebec from 2015–2018). Examination of the participants' durations between tests reveals that the majority of the participants had 335-336 or 371-371 days between assessments - presumably these dates correspond to the timing of the preseason tests for the different sports. Would you have more eligible participants if you had broadened the eligibility criterion? Answer: We thank the reviewers for this comment. Our eligibility criteria only required that the athlete not have a concussion or undergo vision training between tests, and did not have a condition that would affect the results of vision testing. We are not sure which of these criteria the reviewers think we could relax and still obtain an unbiased answer to the question of 1-year test-retest reliability. We could have shortened the interval to only several months, but that would no longer be answering the 1-year test-retest reliability question. We have not made any changes to the manuscript. __________________ Comment: It is unclear how it could be that your participants were limited to waterpolo and short-track speed skating, when presumably you started with a larger number of sports, but this should be clarified as it may reflect a bias in participant selection. As well, both waterpolo (Black et al. 2017) and short-track speed skating (Quinn et al. 2003) have a relatively high rate of concussions, and presumably the athletes may have received subconcussive head impacts, without receiving a concussion. Repetitive hits to the head are associated with microstructural and functional changes in the brain (Mainwaring et al. 2018), and therefore should be acknowledged as a potential factor for the participants in this paper. Answer: We thank the reviewers for raising these points. For the types of sports participants were engaged in, these are the data provided to us. Many athletes from other sports only had 1 test, and some had concussions or vision testing within the 1-year interval. We do not have data on which athletes were referred for testing but never went for the test. The reviewers suggested Black et al reported waterpolo as a sport with many concussions. However, the study cited actually reported 0 concussions in waterpolo athletes. The Quinn et al study reported 6 concussions in 63 athletes over a 1-year period. The authors did not include the injury rate in the paper and it is not possible to compare risks to other sports without knowing how often the athletes were competing / practicing. In general, short-track speed skating concussions occur because of collisions that cause the athlete to fall, and then they may hit their head into the padded boards or on the ice. There are not multiple small hits like one would receive in American football or hockey. That said, we expand on the issue below for other studies that might include athletes from these types of sports. The reviewers suggest cumulative subconcussive head impacts should be raised as potential factor for participants in this study. We are not sure what the reviewers mean. We agree with the paper by Mainwaring et al. (2018) that the reviewers cited. Mainwaring et al states: “Both the research and conceptual understanding of this phenomenon are in their infancy” “the findings are equivocal regarding the effect of subconcussive impacts on the brain” “Insufficient evidence was presented to conclude that repetitive head impacts are associated with neurocognitive impairment. It may be that neuropsychological assessment tools are not sufficiently sensitive to detect any subtle changes in cognitive function that emerge from subconcussive impacts, or that the neurocognitive changes are inconsequential, or follow neurophysiological changes or damage.” “Future research is needed to characterize the phenomenon in question.” As an example, one study found that repetitive subconcussive head impacts over a single season do not appear to result in short-term neurologic impairment (see Gysland SM, Mihalik JP, Register-Mihalik JK, Trulock SC, Shields EW, Guskiewicz KM. The relationship between subconcussive impacts and concussion history on clinical measures of neurologic function in collegiate football players . Annals of biomedical engineering. 2012;40(1):14-22). Aside from these results that do not support a decrease in neurocognitive function with subconcussive impacts, our objective in this study was to report on the 1-year test-retest reliability of vision tests in order to help clinicians understand how to interpret differences between testing conducted post-concussion and at baseline. If subconcussive impacts did affect vision testing, one would expect a decline in visual function as a consequence of subconcussive impacts. If this occurred, any change in test scores between baseline and post-concussion could not be attributed to concussion. That said, we doubt this is the case because if that were true, one would expect a decline in vision function (and test scores) conducted one year after baseline. We did not observe this in our data. However, we acknowledge that the athletes in our study were not involved in sports with many subconcussive impacts. We have modified the text at the end of the limitation, which now includes: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm). In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. Despite starting with athletes from many sports, only athletes from Waterpolo and Short-track speed skating met our eligibility criteria. It is unclear if subconcussion impacts affect neurological function in general 43 . If subconcussion impacts were common in these sports and affected vision testing, we should have seen a systematic decrease in vision capacity between the two tests; this was not observed. Further, if it were present, the effect would be considered part of the “noise” clinicians have to consider when comparing the results from post-concussion and baseline tests.” __________________ Comment: You identify that test-retest reliability of vision tests has been evaluated at the 1-45 days time span. However, studies have evaluated longer-term test-retest reliability. For example, Klein and Fischer (2005) evaluated 19-month test–retest correlations of pro- and anti-saccadic eye movements on 117 participants. Of more direct relevance to the student athletes evaluated in your paper, Breedlove et al. (2019) evaluated the reliability of the King-Devick test (prosaccades) on NCAA athletes, including 833 participants with measures one year apart, and Naidu et al. (2018) evaluated the season-to-season reliability of the King-Devick Test in Canadian professional football players. Your paper would be strengthened by incorporating a fuller complement of relevant papers that have performed longer-term test-retest reliability measures of vision tests, and comparing your findings with theirs. Answer: We thank the reviewers for the reference that we had not been aware of. Our study investigated tests for specific visual function. Although the King-Devick test is sometimes used in concussion, it measures a combination of functions much beyond visual function. Therefore, we do not feel it is relevant to our research questions. We were not aware of the Klein and Fischer article and have now included the reference in the Introduction and Discussion. Our test was quite different from that studied in Klein and Fisher. The introduction text now reads: Previous investigations of the test-retest reliability of these vision tests have used short test-retest time intervals ranging from 1 day to 45 days 9 – 17 , except for one test of saccades 44 . and the Discussion text now reads: “However, we could not find any research examining the stability of the vision tests over a one year period, in athlete or non-athlete populations except for one test of saccades that was very different from the test used in this study 44 .” __________________ Comment: The saccade measures reported in the paper have extremely limited value as they were collected using proprietary equipment - they should likely be removed from the paper. Answer: We respectfully disagree with the reviewers. First, we do not see any harm in including the result of a non-standard test and readers who are not interested can simply ignore the results. Second, we evaluated this non-standard test as this measure was in our a priori protocol. Omitting analyses described in an a priori protocol is a form of reporting bias that we would prefer to avoid. We have modified the limitation section to say: “Finally, the results of the test of Saccades in this study are based on the unpublished proprietary algorithm developed by the clinician. This limits its applicability for other clinicians.” __________________ Comment: The scatterplots (Figures 2A, 3A and 4A) show the line of identity, but it would be interesting to also see the line of best fit. Furthermore, for the parameters with outliers, it would be interesting to add the lines of best fit with and without the outlier. Answer: We thank the reviewers for this comment. We believe the recommended statistical practice for evaluating reliability is the ICC with line of identity, and LOA. We have provided the references that guided this decision. Lines of best fit are not measures of reliability. In addition, any comparison of regression lines with the line of identity can be misleading because one must incorporate the uncertainty due to sampling. If the reviewers have an appropriate statistical reference that supports using regression in studies of test-retest reliability, we would be happy to add the analyses in a subsequent revision. __________________ Comment: The raw data presented through the Data Availability link is very helpful for gaining insight into the specifics of your data. However, it reveals that all of the data are reported as integers. Is this level of precision adequate for capturing the various vision tests? It would be helpful to include a "data dictionary", as recommended for best practices with spreadsheets (Broman and Woo, 2017). Answer: We thank the reviewers for their comment. We have developed a data dictionary and uploaded it as metadata. We agree that some of the measures could have been measured more precisely than others but these are the data provided from the clinician with expertise in orthoptics. We have added text to the beginning of the 2 nd paragraph in the limitations section which now reads: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm).” Competing Interests: No competing interests were disclosed. Close Report a concern Respond or Comment COMMENTS ON THIS REPORT Author Response 26 Aug 2020 Ian Shrier , Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada 26 Aug 2020 Author Response Author Responses REVIEWER # 2&3 Comment: This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. Answer: We thank the ... Continue reading Author Responses REVIEWER # 2&3 Comment: This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. Answer: We thank the reviewers for their interest in our manuscript. __________________ Comment: The outlying data points in Positive Fusional Vergence at 30 cm and 3 m have been identified by the previous external peer reviewers and warrant additional consideration. You present your ICC findings with and without the outlying data points, which is appropriate. However, you do not fully characterize the extreme deviance of the outlying data points. Your wording about the outliers is rather misleading – you state “we noticed one outlier that greatly increased the range of values along x-axis in Figure 2 and Figure 3”. However, both the x- and y-coordinates of the outlier in Figure 2 meet your definition of outlier (1.5 interquartile ranges below the first quartile or above the third quartile), so it is not merely an issue with the x-axis. As well, you state “There was also one outlier for Positive Fusional Vergence at 3 m, 1.5 interquartile range above the third quartile”, but the 1.5 interquartile range (IQR) is your threshold for identifying outliers, not the description of the outlier – this data point is actually 8.125 IQRs above the third quartile. Statistically speaking, it is extremely unlikely that this data point is part of the same distribution as the rest of the data set. Finally, and perhaps most importantly, you state that you “have no reason to believe the data are inaccurate”. However, these outlying data points all exceed the range of normative data for Positive Fusional Vergence that you present in Table 1, providing a strong reason for believing that these data points are questionable. Your previous response “if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point” trivializes the issue. The issues about the outlier data points must be more thoroughly addressed in the manuscript. Answer: We thank the reviewers for raising this point and have modified the text accordingly. For the comment that 1.5 IQR is the threshold and not the data point, we agree and have removed the phrase. It now reads: “There was also one outlier for Positive Fusional Vergence at 3m. When removing this outlier in a sensitivity analysis, the ICC dropped from 0.57 to 0.21.” The paragraph in which this is mentioned refers to the fact that the 1-year test-retest reliability had higher ICC than the 1-week test-retest reliability and this should not be possible. Our sensitivity analyses were conducted to determine if this occurred because of the increased range observed in the 1-year data. The reviewers are correct that the outlier in question is indeed an outlier on both the x and y axis. We have modified the text accordingly. This particular section now reads as below. Similar changes were made to other parts of the manuscript where appropriate: “Given the very high ICC and the presence of an outlier that greatly increased the range of values for the measure (known to increase ICC), we conducted a sensitivity analysis excluding the outlier.” With respect to justifying keeping the outlier in the plot or not, we did not mean to trivialize the issue. We only meant that the decision to remove an outlier needs more justification than simply that the data point was unexpected. Therefore, although we agree with the reviewers that our sensitivity analysis for the ICC is more likely to be correct, we do not feel there is enough evidence to replace the original analysis with the sensitivity analysis as the primary analysis. We feel that discussing this at length would be more confusing than helpful and have deleted the phrase related to “have no reason to believe the data are inaccurate”. The full paragraph now reads: “In one-year test-retest, Positive Fusional Vergence showed excellent reliability at 30cm (ICC=0.93) and moderate at 3m (ICC=0.56), initially. These values were better than the one-week test-retest reliability (ICC=0.54 and 0.49, respectively) 17 . It is difficult to understand how test-retest reliability over one year could be better than test-retest reliability over one week. When we explored the data further, we noticed one outlier that greatly increased the range of values for Positive Fusional Vergence at 30cm (Figure 2) and Positive Fusional Vergence at 3m (Figure 3). Increasing the range of values is known to increase the ICC. This is because ICC is based on the results of an analysis of variance which separates the error into variability between individuals (range of values along x or y axes) and variability within an individual. Therefore, if variability between persons increases, indicated by a larger range of values, ICC will increase. We explored how removing the outlier in our data would affect the results. When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is below the value found for the one-week test-retest reliability; it did not affect LoA. When we removed the outlier (same person) from Positive Fusional Vergence at 3m, the ICC decreased to 0.21. Note that the outlier for this measure had a large difference between the two test scores, and removing such a data point would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” __________________ Comment: It is highly unfortunate that the sample size is so limited, particularly since it would appear that your inclusion criteria were quite broad (followed by the Institut National du Sport du Quebec from 2015–2018). Examination of the participants' durations between tests reveals that the majority of the participants had 335-336 or 371-371 days between assessments - presumably these dates correspond to the timing of the preseason tests for the different sports. Would you have more eligible participants if you had broadened the eligibility criterion? Answer: We thank the reviewers for this comment. Our eligibility criteria only required that the athlete not have a concussion or undergo vision training between tests, and did not have a condition that would affect the results of vision testing. We are not sure which of these criteria the reviewers think we could relax and still obtain an unbiased answer to the question of 1-year test-retest reliability. We could have shortened the interval to only several months, but that would no longer be answering the 1-year test-retest reliability question. We have not made any changes to the manuscript. __________________ Comment: It is unclear how it could be that your participants were limited to waterpolo and short-track speed skating, when presumably you started with a larger number of sports, but this should be clarified as it may reflect a bias in participant selection. As well, both waterpolo (Black et al. 2017) and short-track speed skating (Quinn et al. 2003) have a relatively high rate of concussions, and presumably the athletes may have received subconcussive head impacts, without receiving a concussion. Repetitive hits to the head are associated with microstructural and functional changes in the brain (Mainwaring et al. 2018), and therefore should be acknowledged as a potential factor for the participants in this paper. Answer: We thank the reviewers for raising these points. For the types of sports participants were engaged in, these are the data provided to us. Many athletes from other sports only had 1 test, and some had concussions or vision testing within the 1-year interval. We do not have data on which athletes were referred for testing but never went for the test. The reviewers suggested Black et al reported waterpolo as a sport with many concussions. However, the study cited actually reported 0 concussions in waterpolo athletes. The Quinn et al study reported 6 concussions in 63 athletes over a 1-year period. The authors did not include the injury rate in the paper and it is not possible to compare risks to other sports without knowing how often the athletes were competing / practicing. In general, short-track speed skating concussions occur because of collisions that cause the athlete to fall, and then they may hit their head into the padded boards or on the ice. There are not multiple small hits like one would receive in American football or hockey. That said, we expand on the issue below for other studies that might include athletes from these types of sports. The reviewers suggest cumulative subconcussive head impacts should be raised as potential factor for participants in this study. We are not sure what the reviewers mean. We agree with the paper by Mainwaring et al. (2018) that the reviewers cited. Mainwaring et al states: “Both the research and conceptual understanding of this phenomenon are in their infancy” “the findings are equivocal regarding the effect of subconcussive impacts on the brain” “Insufficient evidence was presented to conclude that repetitive head impacts are associated with neurocognitive impairment. It may be that neuropsychological assessment tools are not sufficiently sensitive to detect any subtle changes in cognitive function that emerge from subconcussive impacts, or that the neurocognitive changes are inconsequential, or follow neurophysiological changes or damage.” “Future research is needed to characterize the phenomenon in question.” As an example, one study found that repetitive subconcussive head impacts over a single season do not appear to result in short-term neurologic impairment (see Gysland SM, Mihalik JP, Register-Mihalik JK, Trulock SC, Shields EW, Guskiewicz KM. The relationship between subconcussive impacts and concussion history on clinical measures of neurologic function in collegiate football players . Annals of biomedical engineering. 2012;40(1):14-22). Aside from these results that do not support a decrease in neurocognitive function with subconcussive impacts, our objective in this study was to report on the 1-year test-retest reliability of vision tests in order to help clinicians understand how to interpret differences between testing conducted post-concussion and at baseline. If subconcussive impacts did affect vision testing, one would expect a decline in visual function as a consequence of subconcussive impacts. If this occurred, any change in test scores between baseline and post-concussion could not be attributed to concussion. That said, we doubt this is the case because if that were true, one would expect a decline in vision function (and test scores) conducted one year after baseline. We did not observe this in our data. However, we acknowledge that the athletes in our study were not involved in sports with many subconcussive impacts. We have modified the text at the end of the limitation, which now includes: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm). In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. Despite starting with athletes from many sports, only athletes from Waterpolo and Short-track speed skating met our eligibility criteria. It is unclear if subconcussion impacts affect neurological function in general 43 . If subconcussion impacts were common in these sports and affected vision testing, we should have seen a systematic decrease in vision capacity between the two tests; this was not observed. Further, if it were present, the effect would be considered part of the “noise” clinicians have to consider when comparing the results from post-concussion and baseline tests.” __________________ Comment: You identify that test-retest reliability of vision tests has been evaluated at the 1-45 days time span. However, studies have evaluated longer-term test-retest reliability. For example, Klein and Fischer (2005) evaluated 19-month test–retest correlations of pro- and anti-saccadic eye movements on 117 participants. Of more direct relevance to the student athletes evaluated in your paper, Breedlove et al. (2019) evaluated the reliability of the King-Devick test (prosaccades) on NCAA athletes, including 833 participants with measures one year apart, and Naidu et al. (2018) evaluated the season-to-season reliability of the King-Devick Test in Canadian professional football players. Your paper would be strengthened by incorporating a fuller complement of relevant papers that have performed longer-term test-retest reliability measures of vision tests, and comparing your findings with theirs. Answer: We thank the reviewers for the reference that we had not been aware of. Our study investigated tests for specific visual function. Although the King-Devick test is sometimes used in concussion, it measures a combination of functions much beyond visual function. Therefore, we do not feel it is relevant to our research questions. We were not aware of the Klein and Fischer article and have now included the reference in the Introduction and Discussion. Our test was quite different from that studied in Klein and Fisher. The introduction text now reads: Previous investigations of the test-retest reliability of these vision tests have used short test-retest time intervals ranging from 1 day to 45 days 9 – 17 , except for one test of saccades 44 . and the Discussion text now reads: “However, we could not find any research examining the stability of the vision tests over a one year period, in athlete or non-athlete populations except for one test of saccades that was very different from the test used in this study 44 .” __________________ Comment: The saccade measures reported in the paper have extremely limited value as they were collected using proprietary equipment - they should likely be removed from the paper. Answer: We respectfully disagree with the reviewers. First, we do not see any harm in including the result of a non-standard test and readers who are not interested can simply ignore the results. Second, we evaluated this non-standard test as this measure was in our a priori protocol. Omitting analyses described in an a priori protocol is a form of reporting bias that we would prefer to avoid. We have modified the limitation section to say: “Finally, the results of the test of Saccades in this study are based on the unpublished proprietary algorithm developed by the clinician. This limits its applicability for other clinicians.” __________________ Comment: The scatterplots (Figures 2A, 3A and 4A) show the line of identity, but it would be interesting to also see the line of best fit. Furthermore, for the parameters with outliers, it would be interesting to add the lines of best fit with and without the outlier. Answer: We thank the reviewers for this comment. We believe the recommended statistical practice for evaluating reliability is the ICC with line of identity, and LOA. We have provided the references that guided this decision. Lines of best fit are not measures of reliability. In addition, any comparison of regression lines with the line of identity can be misleading because one must incorporate the uncertainty due to sampling. If the reviewers have an appropriate statistical reference that supports using regression in studies of test-retest reliability, we would be happy to add the analyses in a subsequent revision. __________________ Comment: The raw data presented through the Data Availability link is very helpful for gaining insight into the specifics of your data. However, it reveals that all of the data are reported as integers. Is this level of precision adequate for capturing the various vision tests? It would be helpful to include a "data dictionary", as recommended for best practices with spreadsheets (Broman and Woo, 2017). Answer: We thank the reviewers for their comment. We have developed a data dictionary and uploaded it as metadata. We agree that some of the measures could have been measured more precisely than others but these are the data provided from the clinician with expertise in orthoptics. We have added text to the beginning of the 2 nd paragraph in the limitations section which now reads: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm).” Author Responses REVIEWER # 2&3 Comment: This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. Answer: We thank the reviewers for their interest in our manuscript. __________________ Comment: The outlying data points in Positive Fusional Vergence at 30 cm and 3 m have been identified by the previous external peer reviewers and warrant additional consideration. You present your ICC findings with and without the outlying data points, which is appropriate. However, you do not fully characterize the extreme deviance of the outlying data points. Your wording about the outliers is rather misleading – you state “we noticed one outlier that greatly increased the range of values along x-axis in Figure 2 and Figure 3”. However, both the x- and y-coordinates of the outlier in Figure 2 meet your definition of outlier (1.5 interquartile ranges below the first quartile or above the third quartile), so it is not merely an issue with the x-axis. As well, you state “There was also one outlier for Positive Fusional Vergence at 3 m, 1.5 interquartile range above the third quartile”, but the 1.5 interquartile range (IQR) is your threshold for identifying outliers, not the description of the outlier – this data point is actually 8.125 IQRs above the third quartile. Statistically speaking, it is extremely unlikely that this data point is part of the same distribution as the rest of the data set. Finally, and perhaps most importantly, you state that you “have no reason to believe the data are inaccurate”. However, these outlying data points all exceed the range of normative data for Positive Fusional Vergence that you present in Table 1, providing a strong reason for believing that these data points are questionable. Your previous response “if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point” trivializes the issue. The issues about the outlier data points must be more thoroughly addressed in the manuscript. Answer: We thank the reviewers for raising this point and have modified the text accordingly. For the comment that 1.5 IQR is the threshold and not the data point, we agree and have removed the phrase. It now reads: “There was also one outlier for Positive Fusional Vergence at 3m. When removing this outlier in a sensitivity analysis, the ICC dropped from 0.57 to 0.21.” The paragraph in which this is mentioned refers to the fact that the 1-year test-retest reliability had higher ICC than the 1-week test-retest reliability and this should not be possible. Our sensitivity analyses were conducted to determine if this occurred because of the increased range observed in the 1-year data. The reviewers are correct that the outlier in question is indeed an outlier on both the x and y axis. We have modified the text accordingly. This particular section now reads as below. Similar changes were made to other parts of the manuscript where appropriate: “Given the very high ICC and the presence of an outlier that greatly increased the range of values for the measure (known to increase ICC), we conducted a sensitivity analysis excluding the outlier.” With respect to justifying keeping the outlier in the plot or not, we did not mean to trivialize the issue. We only meant that the decision to remove an outlier needs more justification than simply that the data point was unexpected. Therefore, although we agree with the reviewers that our sensitivity analysis for the ICC is more likely to be correct, we do not feel there is enough evidence to replace the original analysis with the sensitivity analysis as the primary analysis. We feel that discussing this at length would be more confusing than helpful and have deleted the phrase related to “have no reason to believe the data are inaccurate”. The full paragraph now reads: “In one-year test-retest, Positive Fusional Vergence showed excellent reliability at 30cm (ICC=0.93) and moderate at 3m (ICC=0.56), initially. These values were better than the one-week test-retest reliability (ICC=0.54 and 0.49, respectively) 17 . It is difficult to understand how test-retest reliability over one year could be better than test-retest reliability over one week. When we explored the data further, we noticed one outlier that greatly increased the range of values for Positive Fusional Vergence at 30cm (Figure 2) and Positive Fusional Vergence at 3m (Figure 3). Increasing the range of values is known to increase the ICC. This is because ICC is based on the results of an analysis of variance which separates the error into variability between individuals (range of values along x or y axes) and variability within an individual. Therefore, if variability between persons increases, indicated by a larger range of values, ICC will increase. We explored how removing the outlier in our data would affect the results. When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is below the value found for the one-week test-retest reliability; it did not affect LoA. When we removed the outlier (same person) from Positive Fusional Vergence at 3m, the ICC decreased to 0.21. Note that the outlier for this measure had a large difference between the two test scores, and removing such a data point would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” __________________ Comment: It is highly unfortunate that the sample size is so limited, particularly since it would appear that your inclusion criteria were quite broad (followed by the Institut National du Sport du Quebec from 2015–2018). Examination of the participants' durations between tests reveals that the majority of the participants had 335-336 or 371-371 days between assessments - presumably these dates correspond to the timing of the preseason tests for the different sports. Would you have more eligible participants if you had broadened the eligibility criterion? Answer: We thank the reviewers for this comment. Our eligibility criteria only required that the athlete not have a concussion or undergo vision training between tests, and did not have a condition that would affect the results of vision testing. We are not sure which of these criteria the reviewers think we could relax and still obtain an unbiased answer to the question of 1-year test-retest reliability. We could have shortened the interval to only several months, but that would no longer be answering the 1-year test-retest reliability question. We have not made any changes to the manuscript. __________________ Comment: It is unclear how it could be that your participants were limited to waterpolo and short-track speed skating, when presumably you started with a larger number of sports, but this should be clarified as it may reflect a bias in participant selection. As well, both waterpolo (Black et al. 2017) and short-track speed skating (Quinn et al. 2003) have a relatively high rate of concussions, and presumably the athletes may have received subconcussive head impacts, without receiving a concussion. Repetitive hits to the head are associated with microstructural and functional changes in the brain (Mainwaring et al. 2018), and therefore should be acknowledged as a potential factor for the participants in this paper. Answer: We thank the reviewers for raising these points. For the types of sports participants were engaged in, these are the data provided to us. Many athletes from other sports only had 1 test, and some had concussions or vision testing within the 1-year interval. We do not have data on which athletes were referred for testing but never went for the test. The reviewers suggested Black et al reported waterpolo as a sport with many concussions. However, the study cited actually reported 0 concussions in waterpolo athletes. The Quinn et al study reported 6 concussions in 63 athletes over a 1-year period. The authors did not include the injury rate in the paper and it is not possible to compare risks to other sports without knowing how often the athletes were competing / practicing. In general, short-track speed skating concussions occur because of collisions that cause the athlete to fall, and then they may hit their head into the padded boards or on the ice. There are not multiple small hits like one would receive in American football or hockey. That said, we expand on the issue below for other studies that might include athletes from these types of sports. The reviewers suggest cumulative subconcussive head impacts should be raised as potential factor for participants in this study. We are not sure what the reviewers mean. We agree with the paper by Mainwaring et al. (2018) that the reviewers cited. Mainwaring et al states: “Both the research and conceptual understanding of this phenomenon are in their infancy” “the findings are equivocal regarding the effect of subconcussive impacts on the brain” “Insufficient evidence was presented to conclude that repetitive head impacts are associated with neurocognitive impairment. It may be that neuropsychological assessment tools are not sufficiently sensitive to detect any subtle changes in cognitive function that emerge from subconcussive impacts, or that the neurocognitive changes are inconsequential, or follow neurophysiological changes or damage.” “Future research is needed to characterize the phenomenon in question.” As an example, one study found that repetitive subconcussive head impacts over a single season do not appear to result in short-term neurologic impairment (see Gysland SM, Mihalik JP, Register-Mihalik JK, Trulock SC, Shields EW, Guskiewicz KM. The relationship between subconcussive impacts and concussion history on clinical measures of neurologic function in collegiate football players . Annals of biomedical engineering. 2012;40(1):14-22). Aside from these results that do not support a decrease in neurocognitive function with subconcussive impacts, our objective in this study was to report on the 1-year test-retest reliability of vision tests in order to help clinicians understand how to interpret differences between testing conducted post-concussion and at baseline. If subconcussive impacts did affect vision testing, one would expect a decline in visual function as a consequence of subconcussive impacts. If this occurred, any change in test scores between baseline and post-concussion could not be attributed to concussion. That said, we doubt this is the case because if that were true, one would expect a decline in vision function (and test scores) conducted one year after baseline. We did not observe this in our data. However, we acknowledge that the athletes in our study were not involved in sports with many subconcussive impacts. We have modified the text at the end of the limitation, which now includes: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm). In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. Despite starting with athletes from many sports, only athletes from Waterpolo and Short-track speed skating met our eligibility criteria. It is unclear if subconcussion impacts affect neurological function in general 43 . If subconcussion impacts were common in these sports and affected vision testing, we should have seen a systematic decrease in vision capacity between the two tests; this was not observed. Further, if it were present, the effect would be considered part of the “noise” clinicians have to consider when comparing the results from post-concussion and baseline tests.” __________________ Comment: You identify that test-retest reliability of vision tests has been evaluated at the 1-45 days time span. However, studies have evaluated longer-term test-retest reliability. For example, Klein and Fischer (2005) evaluated 19-month test–retest correlations of pro- and anti-saccadic eye movements on 117 participants. Of more direct relevance to the student athletes evaluated in your paper, Breedlove et al. (2019) evaluated the reliability of the King-Devick test (prosaccades) on NCAA athletes, including 833 participants with measures one year apart, and Naidu et al. (2018) evaluated the season-to-season reliability of the King-Devick Test in Canadian professional football players. Your paper would be strengthened by incorporating a fuller complement of relevant papers that have performed longer-term test-retest reliability measures of vision tests, and comparing your findings with theirs. Answer: We thank the reviewers for the reference that we had not been aware of. Our study investigated tests for specific visual function. Although the King-Devick test is sometimes used in concussion, it measures a combination of functions much beyond visual function. Therefore, we do not feel it is relevant to our research questions. We were not aware of the Klein and Fischer article and have now included the reference in the Introduction and Discussion. Our test was quite different from that studied in Klein and Fisher. The introduction text now reads: Previous investigations of the test-retest reliability of these vision tests have used short test-retest time intervals ranging from 1 day to 45 days 9 – 17 , except for one test of saccades 44 . and the Discussion text now reads: “However, we could not find any research examining the stability of the vision tests over a one year period, in athlete or non-athlete populations except for one test of saccades that was very different from the test used in this study 44 .” __________________ Comment: The saccade measures reported in the paper have extremely limited value as they were collected using proprietary equipment - they should likely be removed from the paper. Answer: We respectfully disagree with the reviewers. First, we do not see any harm in including the result of a non-standard test and readers who are not interested can simply ignore the results. Second, we evaluated this non-standard test as this measure was in our a priori protocol. Omitting analyses described in an a priori protocol is a form of reporting bias that we would prefer to avoid. We have modified the limitation section to say: “Finally, the results of the test of Saccades in this study are based on the unpublished proprietary algorithm developed by the clinician. This limits its applicability for other clinicians.” __________________ Comment: The scatterplots (Figures 2A, 3A and 4A) show the line of identity, but it would be interesting to also see the line of best fit. Furthermore, for the parameters with outliers, it would be interesting to add the lines of best fit with and without the outlier. Answer: We thank the reviewers for this comment. We believe the recommended statistical practice for evaluating reliability is the ICC with line of identity, and LOA. We have provided the references that guided this decision. Lines of best fit are not measures of reliability. In addition, any comparison of regression lines with the line of identity can be misleading because one must incorporate the uncertainty due to sampling. If the reviewers have an appropriate statistical reference that supports using regression in studies of test-retest reliability, we would be happy to add the analyses in a subsequent revision. __________________ Comment: The raw data presented through the Data Availability link is very helpful for gaining insight into the specifics of your data. However, it reveals that all of the data are reported as integers. Is this level of precision adequate for capturing the various vision tests? It would be helpful to include a "data dictionary", as recommended for best practices with spreadsheets (Broman and Woo, 2017). Answer: We thank the reviewers for their comment. We have developed a data dictionary and uploaded it as metadata. We agree that some of the measures could have been measured more precisely than others but these are the data provided from the clinician with expertise in orthoptics. We have added text to the beginning of the 2 nd paragraph in the limitations section which now reads: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm).” Competing Interests: No competing interests were disclosed. Close Report a concern COMMENT ON THIS REPORT Version 1 VERSION 1 PUBLISHED 09 Jul 2019 Views 0 Cite How to cite this report: Haider MN. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.21476.r60656 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v1#referee-response-60656 NOTE: it is important to ensure the information in square brackets after the title is included in this citation. Close Copy Citation Details Reviewer Report 10 Mar 2020 M Nadir Haider , Jacobs School of Medicine and Biomedical Sciences, State University of New York at Buffalo, Buffalo, NY, USA Approved VIEWS 0 https://doi.org/10.5256/f1000research.21476.r60656 Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through ... Continue reading READ ALL Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through the entire manuscript and it is exceptionally well-written, it shows that it has gone through several internal, and even some external, reviews and revisions already. The statistical analysis are correctly described and the appropriate tests and graphs are used to present data. The most obvious downside of this study is the small sample size, there is so much within-subject variation among these test due to the natural process of aging and ocular adaptations which could be due to insignificant events like getting a new monitor for work. Future studies should be performed on larger sample sizes, etc. But I believe that there is merit in having your study indexed for a couple of reasons. The research protocol and analysis are well explained and could be used for design future oculomotor retest reliability studies. Secondly, I am glad that you had concussion as your exclusionary criteria since there are a hundred different publications showing abnormalities in vision function tests after concussion, yet present no retest reliability without the presence of a concussive head injury. I think this paper provides some preliminary evidence which should be made available to other researchers and I think this is a citable manuscript. I do not have any sentence by sentence suggestions, but my only major suggestion is to remove the pre-outlier ICC of Positive Fusional Vergence at 30cm value of 0.93 and say that it is 0.55 (moderate). And I think Negative Fusional Vergence at 30cm should be classified as Good ICC (not moderate since it is between 0.75 and 0.9). Is the work clearly and accurately presented and does it cite the current literature? Yes Is the study design appropriate and is the work technically sound? Yes Are sufficient details of methods and analysis provided to allow replication by others? Yes If applicable, is the statistical analysis and its interpretation appropriate? Yes Are all the source data underlying the results available to ensure full reproducibility? No source data required Are the conclusions drawn adequately supported by the results? Partly Competing Interests: No competing interests were disclosed. Reviewer Expertise: Statistical design, physiological and biochemical markers of concussion, autonomic regulation of cerebral blood blow. I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard. Close READ LESS CITE CITE HOW TO CITE THIS REPORT Haider MN. Reviewer Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.21476.r60656 ) The direct URL for this report is: https://f1000research.com/articles/8-1032/v1#referee-response-60656 NOTE: it is important to ensure the information in square brackets after the title is included in all citations of this article. COPY CITATION DETAILS Report a concern Author Response 31 Mar 2020 Ian Shrier , Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada 31 Mar 2020 Author Response REVIEWER #1 Comment: Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 ... Continue reading REVIEWER #1 Comment: Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through the entire manuscript and it is exceptionally well-written, it shows that it has gone through several internal, and even some external, reviews and revisions already. The statistical analysis are correctly described and the appropriate tests and graphs are used to present data. Answer: We thank the reviewer for the kind comments. ________ Comment: The most obvious downside of this study is the small sample size, there is so much within-subject variation among these test due to the natural process of aging and ocular adaptations which could be due to insignificant events like getting a new monitor for work. Future studies should be performed on larger sample sizes, etc. Answer : In this paper, we used all eligible participants from a clinical database. Therefore, we could not calculate an a priori sample size. Our primary approach to sample size requirements is to estimate precision rather than use hypothesis testing. We have tried to provide information for both approaches in the current version, and the new final paragraph of the limitations section is provided below. "This is a historical cohort observational study, a study design which has inherent limitations. In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. With an effective sample size of 16, the anticipated precision of ICC estimates was +/- 0.25 and the study had 80% power to detect ICC values >= 0.6 and more than 90% power to detect ICC values >=0.7 i.e. rejection of the null hypothesis (Table 1a in 42 ). Note that a total of >60 individuals were required to exclude ICC values 0.7 (Table 2b in 42 )." Reference: Bujang MA, N. B. A simplified guide to determination of sample size requirements for estimating the value of intraclass correlation coefficient: a review. Arch Orofac Sci . 2017; 12(1): 1-11. ______ Comment: But I believe that there is merit in having your study indexed for a couple of reasons. The research protocol and analysis are well explained and could be used for design future oculomotor retest reliability studies. Secondly, I am glad that you had concussion as your exclusionary criteria since there are a hundred different publications showing abnormalities in vision function tests after concussion, yet present no retest reliability without the presence of a concussive head injury. I think this paper provides some preliminary evidence which should be made available to other researchers and I think this is a citable manuscript. Answer : We again thank the reviewer for the kind comments. ________ Comment: I do not have any sentence by sentence suggestions, but my only major suggestion is to remove the pre-outlier ICC of Positive Fusional Vergence at 30cm value of 0.93 and say that it is 0.55 (moderate). Answer: We thank the reviewer for the comment. Recommended practice is to only delete data points if you have a very good reason to believe they are inaccurate. Otherwise, one should keep the original analysis intact and apply sensitivity analyses. For example, if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point. For this reason, we have not changed our results as suggested. However, we have modified the text to further emphasize the importance of the sensitivity analysis. _________ Comment: And I think Negative Fusional Vergence at 30cm should be classified as Good ICC (not moderate since it is between 0.75 and 0.9). Answer: We thank the reviewer for pointing out this oversight. We have now indicated that Figure 3 shows results for good to moderate reliability tests, and made the associated changes in the abstract and manuscript as well. REVIEWER #1 Comment: Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through the entire manuscript and it is exceptionally well-written, it shows that it has gone through several internal, and even some external, reviews and revisions already. The statistical analysis are correctly described and the appropriate tests and graphs are used to present data. Answer: We thank the reviewer for the kind comments. ________ Comment: The most obvious downside of this study is the small sample size, there is so much within-subject variation among these test due to the natural process of aging and ocular adaptations which could be due to insignificant events like getting a new monitor for work. Future studies should be performed on larger sample sizes, etc. Answer : In this paper, we used all eligible participants from a clinical database. Therefore, we could not calculate an a priori sample size. Our primary approach to sample size requirements is to estimate precision rather than use hypothesis testing. We have tried to provide information for both approaches in the current version, and the new final paragraph of the limitations section is provided below. "This is a historical cohort observational study, a study design which has inherent limitations. In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. With an effective sample size of 16, the anticipated precision of ICC estimates was +/- 0.25 and the study had 80% power to detect ICC values >= 0.6 and more than 90% power to detect ICC values >=0.7 i.e. rejection of the null hypothesis (Table 1a in 42 ). Note that a total of >60 individuals were required to exclude ICC values 0.7 (Table 2b in 42 )." Reference: Bujang MA, N. B. A simplified guide to determination of sample size requirements for estimating the value of intraclass correlation coefficient: a review. Arch Orofac Sci . 2017; 12(1): 1-11. ______ Comment: But I believe that there is merit in having your study indexed for a couple of reasons. The research protocol and analysis are well explained and could be used for design future oculomotor retest reliability studies. Secondly, I am glad that you had concussion as your exclusionary criteria since there are a hundred different publications showing abnormalities in vision function tests after concussion, yet present no retest reliability without the presence of a concussive head injury. I think this paper provides some preliminary evidence which should be made available to other researchers and I think this is a citable manuscript. Answer : We again thank the reviewer for the kind comments. ________ Comment: I do not have any sentence by sentence suggestions, but my only major suggestion is to remove the pre-outlier ICC of Positive Fusional Vergence at 30cm value of 0.93 and say that it is 0.55 (moderate). Answer: We thank the reviewer for the comment. Recommended practice is to only delete data points if you have a very good reason to believe they are inaccurate. Otherwise, one should keep the original analysis intact and apply sensitivity analyses. For example, if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point. For this reason, we have not changed our results as suggested. However, we have modified the text to further emphasize the importance of the sensitivity analysis. _________ Comment: And I think Negative Fusional Vergence at 30cm should be classified as Good ICC (not moderate since it is between 0.75 and 0.9). Answer: We thank the reviewer for pointing out this oversight. We have now indicated that Figure 3 shows results for good to moderate reliability tests, and made the associated changes in the abstract and manuscript as well. Competing Interests: No competing interests were disclosed. Close Report a concern Respond or Comment COMMENTS ON THIS REPORT Author Response 31 Mar 2020 Ian Shrier , Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada 31 Mar 2020 Author Response REVIEWER #1 Comment: Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 ... Continue reading REVIEWER #1 Comment: Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through the entire manuscript and it is exceptionally well-written, it shows that it has gone through several internal, and even some external, reviews and revisions already. The statistical analysis are correctly described and the appropriate tests and graphs are used to present data. Answer: We thank the reviewer for the kind comments. ________ Comment: The most obvious downside of this study is the small sample size, there is so much within-subject variation among these test due to the natural process of aging and ocular adaptations which could be due to insignificant events like getting a new monitor for work. Future studies should be performed on larger sample sizes, etc. Answer : In this paper, we used all eligible participants from a clinical database. Therefore, we could not calculate an a priori sample size. Our primary approach to sample size requirements is to estimate precision rather than use hypothesis testing. We have tried to provide information for both approaches in the current version, and the new final paragraph of the limitations section is provided below. "This is a historical cohort observational study, a study design which has inherent limitations. In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. With an effective sample size of 16, the anticipated precision of ICC estimates was +/- 0.25 and the study had 80% power to detect ICC values >= 0.6 and more than 90% power to detect ICC values >=0.7 i.e. rejection of the null hypothesis (Table 1a in 42 ). Note that a total of >60 individuals were required to exclude ICC values 0.7 (Table 2b in 42 )." Reference: Bujang MA, N. B. A simplified guide to determination of sample size requirements for estimating the value of intraclass correlation coefficient: a review. Arch Orofac Sci . 2017; 12(1): 1-11. ______ Comment: But I believe that there is merit in having your study indexed for a couple of reasons. The research protocol and analysis are well explained and could be used for design future oculomotor retest reliability studies. Secondly, I am glad that you had concussion as your exclusionary criteria since there are a hundred different publications showing abnormalities in vision function tests after concussion, yet present no retest reliability without the presence of a concussive head injury. I think this paper provides some preliminary evidence which should be made available to other researchers and I think this is a citable manuscript. Answer : We again thank the reviewer for the kind comments. ________ Comment: I do not have any sentence by sentence suggestions, but my only major suggestion is to remove the pre-outlier ICC of Positive Fusional Vergence at 30cm value of 0.93 and say that it is 0.55 (moderate). Answer: We thank the reviewer for the comment. Recommended practice is to only delete data points if you have a very good reason to believe they are inaccurate. Otherwise, one should keep the original analysis intact and apply sensitivity analyses. For example, if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point. For this reason, we have not changed our results as suggested. However, we have modified the text to further emphasize the importance of the sensitivity analysis. _________ Comment: And I think Negative Fusional Vergence at 30cm should be classified as Good ICC (not moderate since it is between 0.75 and 0.9). Answer: We thank the reviewer for pointing out this oversight. We have now indicated that Figure 3 shows results for good to moderate reliability tests, and made the associated changes in the abstract and manuscript as well. REVIEWER #1 Comment: Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through the entire manuscript and it is exceptionally well-written, it shows that it has gone through several internal, and even some external, reviews and revisions already. The statistical analysis are correctly described and the appropriate tests and graphs are used to present data. Answer: We thank the reviewer for the kind comments. ________ Comment: The most obvious downside of this study is the small sample size, there is so much within-subject variation among these test due to the natural process of aging and ocular adaptations which could be due to insignificant events like getting a new monitor for work. Future studies should be performed on larger sample sizes, etc. Answer : In this paper, we used all eligible participants from a clinical database. Therefore, we could not calculate an a priori sample size. Our primary approach to sample size requirements is to estimate precision rather than use hypothesis testing. We have tried to provide information for both approaches in the current version, and the new final paragraph of the limitations section is provided below. "This is a historical cohort observational study, a study design which has inherent limitations. In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. With an effective sample size of 16, the anticipated precision of ICC estimates was +/- 0.25 and the study had 80% power to detect ICC values >= 0.6 and more than 90% power to detect ICC values >=0.7 i.e. rejection of the null hypothesis (Table 1a in 42 ). Note that a total of >60 individuals were required to exclude ICC values 0.7 (Table 2b in 42 )." Reference: Bujang MA, N. B. A simplified guide to determination of sample size requirements for estimating the value of intraclass correlation coefficient: a review. Arch Orofac Sci . 2017; 12(1): 1-11. ______ Comment: But I believe that there is merit in having your study indexed for a couple of reasons. The research protocol and analysis are well explained and could be used for design future oculomotor retest reliability studies. Secondly, I am glad that you had concussion as your exclusionary criteria since there are a hundred different publications showing abnormalities in vision function tests after concussion, yet present no retest reliability without the presence of a concussive head injury. I think this paper provides some preliminary evidence which should be made available to other researchers and I think this is a citable manuscript. Answer : We again thank the reviewer for the kind comments. ________ Comment: I do not have any sentence by sentence suggestions, but my only major suggestion is to remove the pre-outlier ICC of Positive Fusional Vergence at 30cm value of 0.93 and say that it is 0.55 (moderate). Answer: We thank the reviewer for the comment. Recommended practice is to only delete data points if you have a very good reason to believe they are inaccurate. Otherwise, one should keep the original analysis intact and apply sensitivity analyses. For example, if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point. For this reason, we have not changed our results as suggested. However, we have modified the text to further emphasize the importance of the sensitivity analysis. _________ Comment: And I think Negative Fusional Vergence at 30cm should be classified as Good ICC (not moderate since it is between 0.75 and 0.9). Answer: We thank the reviewer for pointing out this oversight. We have now indicated that Figure 3 shows results for good to moderate reliability tests, and made the associated changes in the abstract and manuscript as well. Competing Interests: No competing interests were disclosed. Close Report a concern COMMENT ON THIS REPORT Comments on this article Comments (0) Version 5 VERSION 5 PUBLISHED 09 Jul 2019 ADD YOUR COMMENT Comment keyboard_arrow_left keyboard_arrow_right Open Peer Review Reviewer Status info_outline Alongside their report, reviewers assign a status to the article: Approved The paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved Fundamental flaws in the paper seriously undermine the findings and conclusions Reviewer Reports Invited Reviewers 1 2 Version 5 (revision) 09 Sep 20 read Version 4 (revision) 26 Aug 20 read read Version 3 (revision) 08 Jun 20 read Version 2 (revision) 31 Mar 20 Version 1 09 Jul 19 read M Nadir Haider , State University of New York at Buffalo, Buffalo, USA Dillon Richards , University of Western Ontario, London, Canada James P Dickey , University of Western Ontario, London, Canada Comments on this article All Comments (0) Add a comment Sign up for content alerts Sign Up You are now signed up to receive this alert Browse by related subjects keyboard_arrow_left Back to all reports Reviewer Report 0 Views copyright © 2020 Dickey J et al. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. 10 Sep 2020 | for Version 5 James P Dickey , School of Kinesiology, University of Western Ontario, London, Ontario, Canada Dillon Richards , Health and Rehabilitation Sciences, University of Western Ontario, London, Canada 0 Views copyright © 2020 Dickey J et al. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. format_quote Cite this report speaker_notes Responses (0) Approved info_outline Alongside their report, reviewers assign a status to the article: Approved The paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved Fundamental flaws in the paper seriously undermine the findings and conclusions The authors have revised the paper to acknowledge "Some data in these athletes appear to be outside the normative range of data previously described for the general population", and have provided access to the raw data through the data availability link. This enables the readers to evaluate the credibility fo the data and interpret the findings accordingly. Competing Interests No competing interests were disclosed. Reviewer Expertise Biomechanics, head impact exposure in sports, concussion We confirm that we have read this submission and believe that we have an appropriate level of expertise to confirm that it is of an acceptable scientific standard. reply Respond to this report Responses (0) Dickey JP and Richards D. Peer Review Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.29392.r71042) NOTE: it is important to ensure the information in square brackets after the title is included in this citation. The direct URL for this report is: https://f1000research.com/articles/8-1032/v5#referee-response-71042 keyboard_arrow_left Back to all reports Reviewer Report 0 Views copyright © 2020 Dickey J et al. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. 02 Sep 2020 | for Version 4 Dillon Richards , Health and Rehabilitation Sciences, University of Western Ontario, London, Canada James P Dickey , School of Kinesiology, University of Western Ontario, London, Ontario, Canada 0 Views copyright © 2020 Dickey J et al. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. format_quote Cite this report speaker_notes Responses (1) Approved With Reservations info_outline Alongside their report, reviewers assign a status to the article: Approved The paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved Fundamental flaws in the paper seriously undermine the findings and conclusions Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points are extreme outliers (four data points at 3 or more IQRs above the third quartile, including one value 8.125 IQRs above the third quartile), and that they exceed the range of normative data for Positive Fusional Vergence that you present in Table 1. Our previous review stated that there is a strong reason for believing that these data points are questionable. In fact there is strong evidence that these data points should be eliminated rather than simply evaluating their influence in a "sensitivity analysis". Competing Interests No competing interests were disclosed. Reviewer Expertise Biomechanics, head impact exposure in sports, concussion. We confirm that we have read this submission and believe that we have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however we have significant reservations, as outlined above. reply Respond to this report Responses (1) Author Response 09 Sep 2020 Ian Shrier, Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada Responses to reviewer 2&3: Comment: Thank you for considering the points that we raised in the review, and we note that your recent revisions better characterize the effects of the outliers. We note that you have not chosen to acknowledge our point that these data points are extreme outliers (four data points at 3 or more IQRs above the third quartile, including one value 8.125 IQRs above the third quartile), and that they exceed the range of normative data for Positive Fusional Vergence that you present in Table 1. Our previous review stated that there is a strong reason for believing that these data points are questionable. In fact there is strong evidence that these data points should be eliminated rather than simply evaluating their influence in a “sensitivity analysis”. Response: We thank the reviewers for their feedback. There are three points raised. The reviewers insist that results excluding outliers be considered the primary analysis, and the results including all the data be considered secondary. The reviewers suggest there are four outliers instead of the 2 outliers we noted. The reviewers suggest that 4 points exceed the range of normative data we provided. 1. Sensitivity Analyses We think that there are different ways to look at outliers and interpret the results. As we mentioned in our previous response to the reviewers, we think that “eliminating” data, as the reviewers suggest, is not the optimal approach. Instead, the sensitivity analysis, as we performed, is a preferable approach unless there is a clear data error. In this way, we are transparent about our data set and our analysis, and the readers can evaluate our findings. F1000Research does not have an editor as an arbitrator when authors and reviewers disagree. Therefore, we have made the changes recommended by the reviewer and the analysis with the excluded data is now considered the primary result. We had already based our conclusions on the analyses after exclusions and have now edited the rest of the text as well. 2. Outliers We are not sure why the reviewer thinks there are four outliers. The data in our study are available online at https://osf.io/gnjdm/ . Here are the calculations for outliers, which we defined using the common standard: 1.5*IQR above the 3 rd quantile. Positive Fusional Vergence 30cm First Test: 25%: 20 75%: 31.25 IQR: 11.25 1.5*IQR: 16.9 Outlier Threshold (75%+1.5*IQR): 48.2 Second Test 25%: 23.75 75%: 30 IQR: 6.25 1.5*IQR: 9.4 Outlier Threshold (75%+1.5*IQR): 39.4 There is only one person with values that should be considered as outliers for positive fusional vergence at 30cm (Id=14). This occurred for both tests (90 on the first test and 85 on the second test). The text now reads: “Given the very high ICC and the presence of an outlier that greatly increased the range of the values for the measure (known to increase ICC), we repeated the analysis excluding the outlier. This decreased the ICC from 0.93 to 0.53, and increased the LoA to ±43.5%. One of the reviewers for this paper has insisted that the analysis without the outlier be considered the primary analysis.” We also added a sentence to the figure legend: “When the analysis was repeated excluding the outlier to the far right, the ICC decreased to 0.53 and the LoA increased to 43.5%.” Positive Fusional Vergence 3m First Test 25%: 17.5 75%: 25 IQR: 7.5 1.5*IQR: 11.3 Outlier Threshold (75%+1.5*IQR): 36.3 Second Test 25%: 17.5 75%: 21.25 IQR: 3.75 1.5*IQR: 5.6 Outlier Threshold (75%+1.5*IQR): 26.8 There are two people with values that should be considered as outliers for positive fusional vergence at 3m (Ids 14 and 15). Particiant 14 is an outlier for both measures, and Particicpant 15 is an outlier for the first test. When Participant 14 was removed the ICC dropped from 0.56 to 0.21 as we reported. If we remove only Participant 15, the ICC actually increases from 0.56 to 0.63. If we remove both outliers, the ICC was 0.45 and the LoA decreased from 60.2% to 41.4%. As per the reviewer’s request, we are reporting the analysis with both outliers removed. The text now reads: “There were two outliers for Positive Fusional Vergence at 3m (one participant on both measures and one participant on only one measure). When we removed both of these outliers, the ICC dropped from 0.56 to 0.45 and the LoA decreased from 60.2% to 41.4%. In both of these cases, the two scores from the outlier were quite different. Although one might anticipate that the ICC would increase by removing such outliers, the ICC actually decreased because the range of values for the measure decreased substantially. As above, one of the reviewers for this paper insisted that the analysis without the outliers be considered the primary analysis.” We have also added a sentence to the figure legend that says: “When the analysis for Positive Fusional Vergence at 3m was repeated excluding the two outliers, the ICC decreased to 0.45 and the LoA decreased to 41.4%.” The results of these tests are also reported in the Discussion. The text in that section now reads: “When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is similar to the value found for the one-week test-retest reliability (ICC=0.54); the LoA increased to 43.5%. When we removed the two outliers from Positive Fusional Vergence at 3m, the ICC decreased to 0.45 and LoA decreased to 41.4%. Note that the outliers for this measure had large differences between the two test scores, and removing such data points would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” 4. Normative Data Range We do not know why some data were outside previously described normative data range. These are the data we received from the clinician doing the test as part of his regular clinical practice. It is possible that previously published normative data for the population does not represent normative data for athletes like those included in our study. We have added one sentence mentioning this in the limitations section of the article. It says: “Some of the data in these athletes appear to be outside the normative range of data previously described for the general population.” View more View less Competing Interests No competing interests were disclosed. reply Respond Report a concern Richards D and Dickey JP. Peer Review Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.28661.r70272) NOTE: it is important to ensure the information in square brackets after the title is included in this citation. The direct URL for this report is: https://f1000research.com/articles/8-1032/v4#referee-response-70272 keyboard_arrow_left Back to all reports Reviewer Report 0 Views copyright © 2020 Haider M. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. 27 Aug 2020 | for Version 4 M Nadir Haider , Jacobs School of Medicine and Biomedical Sciences, State University of New York at Buffalo, Buffalo, NY, USA 0 Views copyright © 2020 Haider M. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. format_quote Cite this report speaker_notes Responses (0) Approved info_outline Alongside their report, reviewers assign a status to the article: Approved The paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved Fundamental flaws in the paper seriously undermine the findings and conclusions No additional comments. Competing Interests No competing interests were disclosed. Reviewer Expertise Concussion, biomarkers, physiology, cerebral blood flow. I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard. reply Respond to this report Responses (0) Haider MN. Peer Review Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.28661.r70273) NOTE: it is important to ensure the information in square brackets after the title is included in this citation. The direct URL for this report is: https://f1000research.com/articles/8-1032/v4#referee-response-70273 keyboard_arrow_left Back to all reports Reviewer Report 0 Views copyright © 2020 Dickey J et al. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. 20 Jul 2020 | for Version 3 Dillon Richards , Health and Rehabilitation Sciences, University of Western Ontario, London, Canada James P Dickey , School of Kinesiology, University of Western Ontario, London, Ontario, Canada 0 Views copyright © 2020 Dickey J et al. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. format_quote Cite this report speaker_notes Responses (1) Approved With Reservations info_outline Alongside their report, reviewers assign a status to the article: Approved The paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved Fundamental flaws in the paper seriously undermine the findings and conclusions This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. The outlying data points in Positive Fusional Vergence at 30 cm and 3 m have been identified by the previous external peer reviewers and warrant additional consideration. You present your ICC findings with and without the outlying data points, which is appropriate. However, you do not fully characterize the extreme deviance of the outlying data points. Your wording about the outliers is rather misleading – you state “we noticed one outlier that greatly increased the range of values along x-axis in Figure 2 and Figure 3”. However, both the x- and y-coordinates of the outlier in Figure 2 meet your definition of outlier (1.5 interquartile ranges below the first quartile or above the third quartile), so it is not merely an issue with the x-axis. As well, you state “There was also one outlier for Positive Fusional Vergence at 3 m, 1.5 interquartile range above the third quartile”, but the 1.5 interquartile range (IQR) is your threshold for identifying outliers, not the description of the outlier – this data point is actually 8.125 IQRs above the third quartile. Statistically speaking, it is extremely unlikely that this data point is part of the same distribution as the rest of the data set. Finally, and perhaps most importantly, you state that you “have no reason to believe the data are inaccurate”. However, these outlying data points all exceed the range of normative data for Positive Fusional Vergence that you present in Table 1, providing a strong reason for believing that these data points are questionable. Your previous response “if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point” trivializes the issue. The issues about the outlier data points must be more thoroughly addressed in the manuscript. It is highly unfortunate that the sample size is so limited, particularly since it would appear that your inclusion criteria were quite broad (followed by the Institut National du Sport du Quebec from 2015–2018). Examination of the participants' durations between tests reveals that the majority of the participants had 335-336 or 371-371 days between assessments - presumably these dates correspond to the timing of the preseason tests for the different sports. Would you have more eligible participants if you had broadened the eligibility criterion? It is unclear how it could be that your participants were limited to waterpolo and short-track speed skating, when presumably you started with a larger number of sports, but this should be clarified as it may reflect a bias in participant selection. As well, both waterpolo (Black et al. 2017) 1 and short-track speed skating (Quinn et al. 2003) 2 have a relatively high rate of concussions, and presumably the athletes may have received subconcussive head impacts, without receiving a concussion. Repetitive hits to the head are associated with microstructural and functional changes in the brain (Mainwaring et al. 2018) 3 , and therefore should be acknowledged as a potential factor for the participants in this paper. You identify that test-retest reliability of vision tests has been evaluated at the 1 day to 45 days time span. However, studies have evaluated longer-term test-retest reliability. For example, Klein and Fischer (2005) 4 evaluated 19-month test–retest correlations of pro- and anti-saccadic eye movements on 117 participants. Of more direct relevance to the student athletes evaluated in your paper, Breedlove et al. (2019) 5 evaluated the reliability of the King-Devick test (prosaccades) on NCAA athletes, including 833 participants with measures one year apart, and Naidu et al. (2018) 6 evaluated the season-to-season reliability of the King-Devick Test in Canadian professional football players. Your paper would be strengthened by incorporating a fuller complement of relevant papers that have performed longer-term test-retest reliability measures of vision tests, and comparing your findings with theirs. The saccade measures reported in the paper have extremely limited value as they were collected using proprietary equipment - they should likely be removed from the paper. The scatterplots (Figures 2A, 3A and 4A) show the line of identity, but it would be interesting to also see the line of best fit. Furthermore, for the parameters with outliers, it would be interesting to add the lines of best fit with and without the outlier. The raw data presented through the Data Availability link is very helpful for gaining insight into the specifics of your data. However, it reveals that all of the data are reported as integers. Is this level of precision adequate for capturing the various vision tests? It would be helpful to include a "data dictionary", as recommended for best practices with spreadsheets (Broman and Woo, 2017) 7 . Is the work clearly and accurately presented and does it cite the current literature? Yes Is the study design appropriate and is the work technically sound? Partly Are sufficient details of methods and analysis provided to allow replication by others? Partly If applicable, is the statistical analysis and its interpretation appropriate? Yes Are all the source data underlying the results available to ensure full reproducibility? Yes Are the conclusions drawn adequately supported by the results? Partly References 1. Black AM, Sergio LE, Macpherson AK: The Epidemiology of Concussions: Number and Nature of Concussions and Time to Recovery Among Female and Male Canadian Varsity Athletes 2008 to 2011. Clin J Sport Med . 2017; 27 (1): 52-56 PubMed Abstract | Publisher Full Text 2. Quinn A, Lun V, McCall J, Overend T: Injuries in short track speed skating. Am J Sports Med . 31 (4): 507-10 PubMed Abstract | Publisher Full Text 3. Mainwaring L, Ferdinand Pennock KM, Mylabathula S, Alavie BZ: Subconcussive head impacts in sport: A systematic review of the evidence. Int J Psychophysiol . 132 (Pt A): 39-54 PubMed Abstract | Publisher Full Text 4. Klein C, Fischer B: Instrumental and test-retest reliability of saccadic measures. Biol Psychol . 2005; 68 (3): 201-13 PubMed Abstract | Publisher Full Text 5. Breedlove KM, Ortega JD, Kaminski TW, Harmon KG, et al.: King-Devick Test Reliability in National Collegiate Athletic Association Athletes: A National Collegiate Athletic Association-Department of Defense Concussion Assessment, Research and Education Report. J Athl Train . 2019; 54 (12): 1241-1246 PubMed Abstract | Publisher Full Text 6. Naidu D, Borza C, Kobitowich T, Mrazik M: Sideline Concussion Assessment: The King-Devick Test in Canadian Professional Football. Journal of Neurotrauma . 2018; 35 (19): 2283-2286 Publisher Full Text 7. Broman K, Woo K: Data Organization in Spreadsheets. The American Statistician . 2018; 72 (1): 2-10 Publisher Full Text Competing Interests No competing interests were disclosed. Reviewer Expertise Biomechanics, head impact exposure in sports, concussion. We confirm that we have read this submission and believe that we have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however we have significant reservations, as outlined above. reply Respond to this report Responses (1) Author Response 26 Aug 2020 Ian Shrier, Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada Author Responses REVIEWER # 2&3 Comment: This is an interesting and important paper based on the prevalence of concussion and the growing appreciation vision tests for diagnosing and assessing concussion. Answer: We thank the reviewers for their interest in our manuscript. __________________ Comment: The outlying data points in Positive Fusional Vergence at 30 cm and 3 m have been identified by the previous external peer reviewers and warrant additional consideration. You present your ICC findings with and without the outlying data points, which is appropriate. However, you do not fully characterize the extreme deviance of the outlying data points. Your wording about the outliers is rather misleading – you state “we noticed one outlier that greatly increased the range of values along x-axis in Figure 2 and Figure 3”. However, both the x- and y-coordinates of the outlier in Figure 2 meet your definition of outlier (1.5 interquartile ranges below the first quartile or above the third quartile), so it is not merely an issue with the x-axis. As well, you state “There was also one outlier for Positive Fusional Vergence at 3 m, 1.5 interquartile range above the third quartile”, but the 1.5 interquartile range (IQR) is your threshold for identifying outliers, not the description of the outlier – this data point is actually 8.125 IQRs above the third quartile. Statistically speaking, it is extremely unlikely that this data point is part of the same distribution as the rest of the data set. Finally, and perhaps most importantly, you state that you “have no reason to believe the data are inaccurate”. However, these outlying data points all exceed the range of normative data for Positive Fusional Vergence that you present in Table 1, providing a strong reason for believing that these data points are questionable. Your previous response “if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point” trivializes the issue. The issues about the outlier data points must be more thoroughly addressed in the manuscript. Answer: We thank the reviewers for raising this point and have modified the text accordingly. For the comment that 1.5 IQR is the threshold and not the data point, we agree and have removed the phrase. It now reads: “There was also one outlier for Positive Fusional Vergence at 3m. When removing this outlier in a sensitivity analysis, the ICC dropped from 0.57 to 0.21.” The paragraph in which this is mentioned refers to the fact that the 1-year test-retest reliability had higher ICC than the 1-week test-retest reliability and this should not be possible. Our sensitivity analyses were conducted to determine if this occurred because of the increased range observed in the 1-year data. The reviewers are correct that the outlier in question is indeed an outlier on both the x and y axis. We have modified the text accordingly. This particular section now reads as below. Similar changes were made to other parts of the manuscript where appropriate: “Given the very high ICC and the presence of an outlier that greatly increased the range of values for the measure (known to increase ICC), we conducted a sensitivity analysis excluding the outlier.” With respect to justifying keeping the outlier in the plot or not, we did not mean to trivialize the issue. We only meant that the decision to remove an outlier needs more justification than simply that the data point was unexpected. Therefore, although we agree with the reviewers that our sensitivity analysis for the ICC is more likely to be correct, we do not feel there is enough evidence to replace the original analysis with the sensitivity analysis as the primary analysis. We feel that discussing this at length would be more confusing than helpful and have deleted the phrase related to “have no reason to believe the data are inaccurate”. The full paragraph now reads: “In one-year test-retest, Positive Fusional Vergence showed excellent reliability at 30cm (ICC=0.93) and moderate at 3m (ICC=0.56), initially. These values were better than the one-week test-retest reliability (ICC=0.54 and 0.49, respectively) 17 . It is difficult to understand how test-retest reliability over one year could be better than test-retest reliability over one week. When we explored the data further, we noticed one outlier that greatly increased the range of values for Positive Fusional Vergence at 30cm (Figure 2) and Positive Fusional Vergence at 3m (Figure 3). Increasing the range of values is known to increase the ICC. This is because ICC is based on the results of an analysis of variance which separates the error into variability between individuals (range of values along x or y axes) and variability within an individual. Therefore, if variability between persons increases, indicated by a larger range of values, ICC will increase. We explored how removing the outlier in our data would affect the results. When we removed the outlier for Positive Fusional Vergence at 30cm, the ICC dropped to 0.53, which is below the value found for the one-week test-retest reliability; it did not affect LoA. When we removed the outlier (same person) from Positive Fusional Vergence at 3m, the ICC decreased to 0.21. Note that the outlier for this measure had a large difference between the two test scores, and removing such a data point would normally be expected to increase the ICC ( Figure 3). The finding that the ICC decreased indicates that as expected, if the range of values among the populations is similar, the one-year test-retest reliability for Positive Fusional Vergence at both 30cm and 3m is likely less than the one-week test-retest reliability.” __________________ Comment: It is highly unfortunate that the sample size is so limited, particularly since it would appear that your inclusion criteria were quite broad (followed by the Institut National du Sport du Quebec from 2015–2018). Examination of the participants' durations between tests reveals that the majority of the participants had 335-336 or 371-371 days between assessments - presumably these dates correspond to the timing of the preseason tests for the different sports. Would you have more eligible participants if you had broadened the eligibility criterion? Answer: We thank the reviewers for this comment. Our eligibility criteria only required that the athlete not have a concussion or undergo vision training between tests, and did not have a condition that would affect the results of vision testing. We are not sure which of these criteria the reviewers think we could relax and still obtain an unbiased answer to the question of 1-year test-retest reliability. We could have shortened the interval to only several months, but that would no longer be answering the 1-year test-retest reliability question. We have not made any changes to the manuscript. __________________ Comment: It is unclear how it could be that your participants were limited to waterpolo and short-track speed skating, when presumably you started with a larger number of sports, but this should be clarified as it may reflect a bias in participant selection. As well, both waterpolo (Black et al. 2017) and short-track speed skating (Quinn et al. 2003) have a relatively high rate of concussions, and presumably the athletes may have received subconcussive head impacts, without receiving a concussion. Repetitive hits to the head are associated with microstructural and functional changes in the brain (Mainwaring et al. 2018), and therefore should be acknowledged as a potential factor for the participants in this paper. Answer: We thank the reviewers for raising these points. For the types of sports participants were engaged in, these are the data provided to us. Many athletes from other sports only had 1 test, and some had concussions or vision testing within the 1-year interval. We do not have data on which athletes were referred for testing but never went for the test. The reviewers suggested Black et al reported waterpolo as a sport with many concussions. However, the study cited actually reported 0 concussions in waterpolo athletes. The Quinn et al study reported 6 concussions in 63 athletes over a 1-year period. The authors did not include the injury rate in the paper and it is not possible to compare risks to other sports without knowing how often the athletes were competing / practicing. In general, short-track speed skating concussions occur because of collisions that cause the athlete to fall, and then they may hit their head into the padded boards or on the ice. There are not multiple small hits like one would receive in American football or hockey. That said, we expand on the issue below for other studies that might include athletes from these types of sports. The reviewers suggest cumulative subconcussive head impacts should be raised as potential factor for participants in this study. We are not sure what the reviewers mean. We agree with the paper by Mainwaring et al. (2018) that the reviewers cited. Mainwaring et al states: “Both the research and conceptual understanding of this phenomenon are in their infancy” “the findings are equivocal regarding the effect of subconcussive impacts on the brain” “Insufficient evidence was presented to conclude that repetitive head impacts are associated with neurocognitive impairment. It may be that neuropsychological assessment tools are not sufficiently sensitive to detect any subtle changes in cognitive function that emerge from subconcussive impacts, or that the neurocognitive changes are inconsequential, or follow neurophysiological changes or damage.” “Future research is needed to characterize the phenomenon in question.” As an example, one study found that repetitive subconcussive head impacts over a single season do not appear to result in short-term neurologic impairment (see Gysland SM, Mihalik JP, Register-Mihalik JK, Trulock SC, Shields EW, Guskiewicz KM. The relationship between subconcussive impacts and concussion history on clinical measures of neurologic function in collegiate football players . Annals of biomedical engineering. 2012;40(1):14-22). Aside from these results that do not support a decrease in neurocognitive function with subconcussive impacts, our objective in this study was to report on the 1-year test-retest reliability of vision tests in order to help clinicians understand how to interpret differences between testing conducted post-concussion and at baseline. If subconcussive impacts did affect vision testing, one would expect a decline in visual function as a consequence of subconcussive impacts. If this occurred, any change in test scores between baseline and post-concussion could not be attributed to concussion. That said, we doubt this is the case because if that were true, one would expect a decline in vision function (and test scores) conducted one year after baseline. We did not observe this in our data. However, we acknowledge that the athletes in our study were not involved in sports with many subconcussive impacts. We have modified the text at the end of the limitation, which now includes: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm). In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. Despite starting with athletes from many sports, only athletes from Waterpolo and Short-track speed skating met our eligibility criteria. It is unclear if subconcussion impacts affect neurological function in general 43 . If subconcussion impacts were common in these sports and affected vision testing, we should have seen a systematic decrease in vision capacity between the two tests; this was not observed. Further, if it were present, the effect would be considered part of the “noise” clinicians have to consider when comparing the results from post-concussion and baseline tests.” __________________ Comment: You identify that test-retest reliability of vision tests has been evaluated at the 1-45 days time span. However, studies have evaluated longer-term test-retest reliability. For example, Klein and Fischer (2005) evaluated 19-month test–retest correlations of pro- and anti-saccadic eye movements on 117 participants. Of more direct relevance to the student athletes evaluated in your paper, Breedlove et al. (2019) evaluated the reliability of the King-Devick test (prosaccades) on NCAA athletes, including 833 participants with measures one year apart, and Naidu et al. (2018) evaluated the season-to-season reliability of the King-Devick Test in Canadian professional football players. Your paper would be strengthened by incorporating a fuller complement of relevant papers that have performed longer-term test-retest reliability measures of vision tests, and comparing your findings with theirs. Answer: We thank the reviewers for the reference that we had not been aware of. Our study investigated tests for specific visual function. Although the King-Devick test is sometimes used in concussion, it measures a combination of functions much beyond visual function. Therefore, we do not feel it is relevant to our research questions. We were not aware of the Klein and Fischer article and have now included the reference in the Introduction and Discussion. Our test was quite different from that studied in Klein and Fisher. The introduction text now reads: Previous investigations of the test-retest reliability of these vision tests have used short test-retest time intervals ranging from 1 day to 45 days 9 – 17 , except for one test of saccades 44 . and the Discussion text now reads: “However, we could not find any research examining the stability of the vision tests over a one year period, in athlete or non-athlete populations except for one test of saccades that was very different from the test used in this study 44 .” __________________ Comment: The saccade measures reported in the paper have extremely limited value as they were collected using proprietary equipment - they should likely be removed from the paper. Answer: We respectfully disagree with the reviewers. First, we do not see any harm in including the result of a non-standard test and readers who are not interested can simply ignore the results. Second, we evaluated this non-standard test as this measure was in our a priori protocol. Omitting analyses described in an a priori protocol is a form of reporting bias that we would prefer to avoid. We have modified the limitation section to say: “Finally, the results of the test of Saccades in this study are based on the unpublished proprietary algorithm developed by the clinician. This limits its applicability for other clinicians.” __________________ Comment: The scatterplots (Figures 2A, 3A and 4A) show the line of identity, but it would be interesting to also see the line of best fit. Furthermore, for the parameters with outliers, it would be interesting to add the lines of best fit with and without the outlier. Answer: We thank the reviewers for this comment. We believe the recommended statistical practice for evaluating reliability is the ICC with line of identity, and LOA. We have provided the references that guided this decision. Lines of best fit are not measures of reliability. In addition, any comparison of regression lines with the line of identity can be misleading because one must incorporate the uncertainty due to sampling. If the reviewers have an appropriate statistical reference that supports using regression in studies of test-retest reliability, we would be happy to add the analyses in a subsequent revision. __________________ Comment: The raw data presented through the Data Availability link is very helpful for gaining insight into the specifics of your data. However, it reveals that all of the data are reported as integers. Is this level of precision adequate for capturing the various vision tests? It would be helpful to include a "data dictionary", as recommended for best practices with spreadsheets (Broman and Woo, 2017). Answer: We thank the reviewers for their comment. We have developed a data dictionary and uploaded it as metadata. We agree that some of the measures could have been measured more precisely than others but these are the data provided from the clinician with expertise in orthoptics. We have added text to the beginning of the 2 nd paragraph in the limitations section which now reads: “This is a historical cohort observational study, a study design which has inherent limitations. The data provided were not always as precise as one might expect (e.g. near point convergence measured to the nearest cm).” View more View less Competing Interests No competing interests were disclosed. reply Respond Report a concern Richards D and Dickey JP. Peer Review Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.27076.r66529) NOTE: it is important to ensure the information in square brackets after the title is included in this citation. The direct URL for this report is: https://f1000research.com/articles/8-1032/v3#referee-response-66529 keyboard_arrow_left Back to all reports Reviewer Report 0 Views copyright © 2020 Haider M. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. 10 Mar 2020 | for Version 1 M Nadir Haider , Jacobs School of Medicine and Biomedical Sciences, State University of New York at Buffalo, Buffalo, NY, USA 0 Views copyright © 2020 Haider M. This is an open access peer review report distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. format_quote Cite this report speaker_notes Responses (1) Approved info_outline Alongside their report, reviewers assign a status to the article: Approved The paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved Fundamental flaws in the paper seriously undermine the findings and conclusions Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through the entire manuscript and it is exceptionally well-written, it shows that it has gone through several internal, and even some external, reviews and revisions already. The statistical analysis are correctly described and the appropriate tests and graphs are used to present data. The most obvious downside of this study is the small sample size, there is so much within-subject variation among these test due to the natural process of aging and ocular adaptations which could be due to insignificant events like getting a new monitor for work. Future studies should be performed on larger sample sizes, etc. But I believe that there is merit in having your study indexed for a couple of reasons. The research protocol and analysis are well explained and could be used for design future oculomotor retest reliability studies. Secondly, I am glad that you had concussion as your exclusionary criteria since there are a hundred different publications showing abnormalities in vision function tests after concussion, yet present no retest reliability without the presence of a concussive head injury. I think this paper provides some preliminary evidence which should be made available to other researchers and I think this is a citable manuscript. I do not have any sentence by sentence suggestions, but my only major suggestion is to remove the pre-outlier ICC of Positive Fusional Vergence at 30cm value of 0.93 and say that it is 0.55 (moderate). And I think Negative Fusional Vergence at 30cm should be classified as Good ICC (not moderate since it is between 0.75 and 0.9). Is the work clearly and accurately presented and does it cite the current literature? Yes Is the study design appropriate and is the work technically sound? Yes Are sufficient details of methods and analysis provided to allow replication by others? Yes If applicable, is the statistical analysis and its interpretation appropriate? Yes Are all the source data underlying the results available to ensure full reproducibility? No source data required Are the conclusions drawn adequately supported by the results? Partly Competing Interests No competing interests were disclosed. Reviewer Expertise Statistical design, physiological and biochemical markers of concussion, autonomic regulation of cerebral blood blow. I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard. reply Respond to this report Responses (1) Author Response 31 Mar 2020 Ian Shrier, Centre for Clinical Epidemiology, Lady Davis Institute, Jewish General Hospital, McGill University, Montreal, Canada REVIEWER #1 Comment: Thank you for giving me the opportunity to review this manuscript. It measures the retest reliability of common ocular/oculomotor tests over one year. The sample size is 16 college-aged athletes. Intra-class correlation is performed and presented. I have read through the entire manuscript and it is exceptionally well-written, it shows that it has gone through several internal, and even some external, reviews and revisions already. The statistical analysis are correctly described and the appropriate tests and graphs are used to present data. Answer: We thank the reviewer for the kind comments. ________ Comment: The most obvious downside of this study is the small sample size, there is so much within-subject variation among these test due to the natural process of aging and ocular adaptations which could be due to insignificant events like getting a new monitor for work. Future studies should be performed on larger sample sizes, etc. Answer : In this paper, we used all eligible participants from a clinical database. Therefore, we could not calculate an a priori sample size. Our primary approach to sample size requirements is to estimate precision rather than use hypothesis testing. We have tried to provide information for both approaches in the current version, and the new final paragraph of the limitations section is provided below. "This is a historical cohort observational study, a study design which has inherent limitations. In addition, the sample size was relatively small and composed of healthy athletes, which will limit the generalizability of these findings to other populations. Although we started with a pool of 199 athletes, many athletes were excluded because they only had one baseline test, a concussion occurred in between the two baseline tests, or the second baseline test occurred outside the testing window of 365±30 days. With an effective sample size of 16, the anticipated precision of ICC estimates was +/- 0.25 and the study had 80% power to detect ICC values >= 0.6 and more than 90% power to detect ICC values >=0.7 i.e. rejection of the null hypothesis (Table 1a in 42 ). Note that a total of >60 individuals were required to exclude ICC values 0.7 (Table 2b in 42 )." Reference: Bujang MA, N. B. A simplified guide to determination of sample size requirements for estimating the value of intraclass correlation coefficient: a review. Arch Orofac Sci . 2017; 12(1): 1-11. ______ Comment: But I believe that there is merit in having your study indexed for a couple of reasons. The research protocol and analysis are well explained and could be used for design future oculomotor retest reliability studies. Secondly, I am glad that you had concussion as your exclusionary criteria since there are a hundred different publications showing abnormalities in vision function tests after concussion, yet present no retest reliability without the presence of a concussive head injury. I think this paper provides some preliminary evidence which should be made available to other researchers and I think this is a citable manuscript. Answer : We again thank the reviewer for the kind comments. ________ Comment: I do not have any sentence by sentence suggestions, but my only major suggestion is to remove the pre-outlier ICC of Positive Fusional Vergence at 30cm value of 0.93 and say that it is 0.55 (moderate). Answer: We thank the reviewer for the comment. Recommended practice is to only delete data points if you have a very good reason to believe they are inaccurate. Otherwise, one should keep the original analysis intact and apply sensitivity analyses. For example, if deleting a point improved the ICC, we are confident the reviewer would agree that we should not delete the data point. For this reason, we have not changed our results as suggested. However, we have modified the text to further emphasize the importance of the sensitivity analysis. _________ Comment: And I think Negative Fusional Vergence at 30cm should be classified as Good ICC (not moderate since it is between 0.75 and 0.9). Answer: We thank the reviewer for pointing out this oversight. We have now indicated that Figure 3 shows results for good to moderate reliability tests, and made the associated changes in the abstract and manuscript as well. View more View less Competing Interests No competing interests were disclosed. reply Respond Report a concern Haider MN. Peer Review Report For: One-year test-retest reliability of ten vision tests in Canadian athletes [version 5; peer review: 2 approved] . F1000Research 2020, 8 :1032 ( https://doi.org/10.5256/f1000research.21476.r60656) NOTE: it is important to ensure the information in square brackets after the title is included in this citation. The direct URL for this report is: https://f1000research.com/articles/8-1032/v1#referee-response-60656 Alongside their report, reviewers assign a status to the article: Approved - the paper is scientifically sound in its current form and only minor, if any, improvements are suggested Approved with reservations - A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit. Not approved - fundamental flaws in the paper seriously undermine the findings and conclusions Adjust parameters to alter display View on desktop for interactive features Includes Interactive Elements View on desktop for interactive features Competing Interests Policy Provide sufficient details of any financial or non-financial competing interests to enable users to assess whether your comments might lead a reasonable person to question your impartiality. Consider the following examples, but note that this is not an exhaustive list: Examples of 'Non-Financial Competing Interests' Within the past 4 years, you have held joint grants, published or collaborated with any of the authors of the selected paper. You have a close personal relationship (e.g. parent, spouse, sibling, or domestic partner) with any of the authors. You are a close professional associate of any of the authors (e.g. scientific mentor, recent student). You work at the same institute as any of the authors. You hope/expect to benefit (e.g. favour or employment) as a result of your submission. You are an Editor for the journal in which the article is published. Examples of 'Financial Competing Interests' You expect to receive, or in the past 4 years have received, any of the following from any commercial organisation that may gain financially from your submission: a salary, fees, funding, reimbursements. You expect to receive, or in the past 4 years have received, shared grant support or other funding with any of the authors. You hold, or are currently applying for, any patents or significant stocks/shares relating to the subject matter of the paper you are commenting on. Stay Updated Sign up for content alerts and receive a weekly or monthly email with all newly published articles Register with F1000Research Already registered? Sign in Not now, thanks close PLEASE NOTE If you are an AUTHOR of this article, please check that you signed in with the account associated with this article otherwise we cannot automatically identify your role as an author and your comment will be labelled as a “User Comment”. If you are a REVIEWER of this article, please check that you have signed in with the account associated with this article and then go to your account to submit your report, please do not post your review here. If you do not have access to your original account, please contact us . All commenters must hold a formal affiliation as per our Policies . The information that you give us will be displayed next to your comment. User comments must be in English, comprehensible and relevant to the article under discussion. We reserve the right to remove any comments that we consider to be inappropriate, offensive or otherwise in breach of the User Comment Terms and Conditions . Commenters must not use a comment for personal attacks. When criticisms of the article are based on unpublished data, the data should be made available. I accept the User Comment Terms and Conditions Please confirm that you accept the User Comment Terms and Conditions. Affiliation ✕ refresh Please enter your institution. Note: To add your institution or organisation, start typing the name and then select the correct name from the list. Where applicable, the name will appear in both the original language and in English. Do not paste in the name. If the name does not appear in the drop-down list, we will display the information you have entered. ✕ refresh Country/Region * USA UK Canada China France Germany Afghanistan Aland Islands Albania Algeria American Samoa Andorra Angola Anguilla Antarctica Antigua and Barbuda Argentina Armenia Aruba Australia Austria Azerbaijan Bahamas Bahrain Bangladesh Barbados Belarus Belgium Belize Benin Bermuda Bhutan Bolivia Bosnia and Herzegovina Botswana Bouvet Island Brazil British Indian Ocean Territory British Virgin Islands Brunei Bulgaria Burkina Faso Burundi Cambodia Cameroon Canada Cape Verde Cayman Islands Central African Republic Chad Chile China Christmas Island Cocos (Keeling) Islands Colombia Comoros Congo Cook Islands Costa Rica Cote d'Ivoire Croatia Cuba Cyprus Czech Republic Democratic Republic of the Congo Denmark Djibouti Dominica Dominican Republic Ecuador Egypt El Salvador Equatorial Guinea Eritrea Estonia Ethiopia Falkland Islands Faroe Islands Federated States of Micronesia Fiji Finland France French Guiana French Polynesia French Southern Territories Gabon Georgia Germany Ghana Gibraltar Greece Greenland Grenada Guadeloupe Guam Guatemala Guernsey Guinea Guinea-Bissau Guyana Haiti Heard Island and Mcdonald Islands Holy See (Vatican City State) Honduras Hong Kong Hungary Iceland India Indonesia Iran Iraq Ireland Israel Italy Jamaica Japan Jersey Jordan Kazakhstan Kenya Kiribati Kosovo (Serbia and Montenegro) Kuwait Kyrgyzstan Lao People's Democratic Republic Latvia Lebanon Lesotho Liberia Libya Liechtenstein Lithuania Luxembourg Macao Madagascar Malawi Malaysia Maldives Mali Malta Marshall Islands Martinique Mauritania Mauritius Mayotte Mexico Minor Outlying Islands of the United States Moldova Monaco Mongolia Montenegro Montserrat Morocco Mozambique Myanmar Namibia Nauru Nepal Netherlands Antilles New Caledonia New Zealand Nicaragua Niger Nigeria Niue Norfolk Island North Korea North Macedonia Northern Mariana Islands Norway Oman Pakistan Palau Palestinian Territory Panama Papua New Guinea Paraguay Peru Philippines Pitcairn Poland Portugal Puerto Rico Qatar Reunion Romania Russian Federation Rwanda Saint Helena Saint Kitts and Nevis Saint Lucia Saint Pierre and Miquelon Saint Vincent and the Grenadines Samoa San Marino Sao Tome and Principe Saudi Arabia Senegal Serbia Seychelles Sierra Leone Singapore Slovakia Slovenia Solomon Islands Somalia South Africa South Georgia and the South Sandwich Is South Korea South Sudan Spain Sri Lanka Sudan Suriname Svalbard and Jan Mayen Swaziland Sweden Switzerland Syria Taiwan Tajikistan Tanzania Thailand The Gambia The Netherlands Timor-Leste Togo Tokelau Tonga Trinidad and Tobago Tunisia Turkey Turkmenistan Turks and Caicos Islands Tuvalu UK USA Uganda Ukraine United Arab Emirates United States Virgin Islands Uruguay Uzbekistan Vanuatu Venezuela Vietnam Wallis and Futuna West Bank and Gaza Strip Western Sahara Yemen Zambia Zimbabwe Please select your country/region. You must enter a comment. Competing Interests Please disclose any competing interests that might be construed to influence your judgment of the article's or peer review report's validity or importance. Competing Interests Policy Provide sufficient details of any financial or non-financial competing interests to enable users to assess whether your comments might lead a reasonable person to question your impartiality. Consider the following examples, but note that this is not an exhaustive list: Examples of 'Non-Financial Competing Interests' Within the past 4 years, you have held joint grants, published or collaborated with any of the authors of the selected paper. You have a close personal relationship (e.g. parent, spouse, sibling, or domestic partner) with any of the authors. You are a close professional associate of any of the authors (e.g. scientific mentor, recent student). You work at the same institute as any of the authors. You hope/expect to benefit (e.g. favour or employment) as a result of your submission. You are an Editor for the journal in which the article is published. Examples of 'Financial Competing Interests' You expect to receive, or in the past 4 years have received, any of the following from any commercial organisation that may gain financially from your submission: a salary, fees, funding, reimbursements. You expect to receive, or in the past 4 years have received, shared grant support or other funding with any of the authors. You hold, or are currently applying for, any patents or significant stocks/shares relating to the subject matter of the paper you are commenting on. Please state your competing interests The comment has been saved. An error has occurred. Please try again. Cancel Post var lTitle = "One-year test-retest reliability of ten vision...".replace("'", ''); var linkedInUrl = "http://www.linkedin.com/shareArticle?url=https://f1000research.com/articles/8-1032/v5" + "&title=" + encodeURIComponent(lTitle) + "&summary=" + encodeURIComponent('Read the article by '); var deliciousUrl = "https://del.icio.us/post?url=https://f1000research.com/articles/8-1032/v5&title=" + encodeURIComponent(lTitle); var redditUrl = "http://reddit.com/submit?url=https://f1000research.com/articles/8-1032/v5" + "&title=" + encodeURIComponent(lTitle); linkedInUrl += encodeURIComponent('Aloosh M et al.'); var offsetTop = /chrome/i.test( navigator.userAgent ) ? 4 : -10; var addthis_config = { ui_offset_top: offsetTop, services_compact : "facebook,twitter,www.linkedin.com,www.mendeley.com,reddit.com", services_expanded : "facebook,twitter,www.linkedin.com,www.mendeley.com,reddit.com", services_custom : [ { name: "LinkedIn", url: linkedInUrl, icon:"/img/icon/at_linkedin.svg" }, { name: "Mendeley", url: "http://www.mendeley.com/import/?url=https://f1000research.com/articles/8-1032/v5/mendeley", icon:"/img/icon/at_mendeley.svg" }, { name: "Reddit", url: redditUrl, icon:"/img/icon/at_reddit.svg" }, ] }; var addthis_share = { url: "https://f1000research.com/articles/8-1032", templates : { twitter : "One-year test-retest reliability of ten vision tests in Canadian.... Aloosh M et al., published by " + "@F1000Research" + ", https://f1000research.com/articles/8-1032/v5" } }; if (typeof(addthis) != "undefined"){ addthis.addEventListener('addthis.ready', checkCount); addthis.addEventListener('addthis.menu.share', checkCount); } $(".f1r-shares-twitter").attr("href", "https://twitter.com/intent/tweet?text=" + addthis_share.templates.twitter); $(".f1r-shares-facebook").attr("href", "https://www.facebook.com/sharer/sharer.php?u=" + addthis_share.url); $(".f1r-shares-linkedin").attr("href", addthis_config.services_custom[0].url); $(".f1r-shares-reddit").attr("href", addthis_config.services_custom[2].url); $(".f1r-shares-mendelay").attr("href", addthis_config.services_custom[1].url); function checkCount(){ setTimeout(function(){ $(".addthis_button_expanded").each(function(){ var count = $(this).text(); if (count !== "" && count != "0") $(this).removeClass("is-hidden"); else $(this).addClass("is-hidden"); }); }, 1000); } close How to cite this report {{reportCitation}} Cancel Copy Citation Details $(function(){R.ui.buttonDropdowns('.dropdown-for-downloads');}); $(function(){R.ui.toolbarDropdowns('.toolbar-dropdown-for-downloads');}); $.get("/articles/acj/19587/29392") new F1000.Clipboard(); new F1000.ThesaurusTermsDisplay("articles", "article", "29392"); $(document).ready(function() { $( "#frame1" ).on('load', function() { var mydiv = $(this).contents().find("div"); var h = mydiv.height(); console.log(h) }); var tooltipLivingFigure = jQuery(".interactive-living-figure-label .icon-more-info"), titleLivingFigure = tooltipLivingFigure.attr("title"); tooltipLivingFigure.simpletip({ fixed: true, position: ["-115", "30"], baseClass: 'small-tooltip', content:titleLivingFigure + " " }); tooltipLivingFigure.removeAttr("title"); $("body").on("click", ".cite-living-figure", function(e) { e.preventDefault(); var ref = $(this).attr("data-ref"); $(this).closest(".living-figure-list-container").find("#" + ref).fadeIn(200); }); $("body").on("click", ".close-cite-living-figure", function(e) { e.preventDefault(); $(this).closest(".popup-window-wrapper").fadeOut(200); }); $(document).on("mouseup", function(e) { var metricsContainer = $(".article-metrics-popover-wrapper"); if (!metricsContainer.is(e.target) && metricsContainer.has(e.target).length === 0) { $(".article-metrics-close-button").click(); } }); var articleId = $('#articleId').val(); if($("#main-article-count-box").attachArticleMetrics) { $("#main-article-count-box").attachArticleMetrics(articleId, { articleMetricsView: true }); } }); var figshareWidget = $(".new_figshare_widget"); if (figshareWidget.length > 0) { window.figshare.load("f1000", function(Widget) { // Select a tag/tags defined in your page. In this tag we will place the widget. _.map(figshareWidget, function(el){ var widget = new Widget({ articleId: $(el).attr("figshare_articleId") //height:300 // this is the height of the viewer part. [Default: 550] }); widget.initialize(); // initialize the widget widget.mount(el); // mount it in a tag that's on your page // this will save the widget on the global scope for later use from // your JS scripts. This line is optional. //window.widget = widget; }); }); } close Error Close Add Reset F1000.MICROSERVICES.AFFILIATION = ''; $(document).ready(function () { $('.js-affiliations-form').each((index, form) => { new AffiliationForm({ formId: form.id, institutionErrorSelector: '.comment-enter-institution', departmentErrorSelector: '.comment-enter-department', placeSelector: '.js-add-comment-place', stateSelector: '.js-add-comment-state', zipCodeSelector: '.js-add-comment-zipcode', countrySelector: '.js-add-comment-country', countryErrorSelector: '.comment-enter-country', }); }); }); $(document).ready(function () { var reportIds = { "71041": 0, "70273": 14, "70272": 21, "71042": 15, "61838": 0, "64407": 0, "51615": 0, "50975": 0, "50976": 0, "50977": 0, "50978": 0, "50979": 0, "64811": 0, "54068": 0, "54069": 0, "54070": 0, "54071": 0, "54072": 0, "55992": 0, "55993": 0, "55994": 0, "55995": 0, "64828": 0, "64829": 0, "64830": 0, "52287": 0, "52288": 0, "52289": 0, "52290": 0, "51652": 0, "53188": 0, "51653": 0, "53189": 0, "51654": 0, "53190": 0, "51655": 0, "53191": 0, "53192": 0, "54730": 0, "57035": 0, "54731": 0, "54732": 0, "57036": 0, "54733": 0, "57037": 0, "57038": 0, "53082": 0, "66526": 0, "66529": 30, "66528": 0, "63461": 0, "66532": 0, "63462": 0, "63463": 0, "60654": 0, "60655": 0, "60656": 29, "60657": 0, }; $(".referee-response-container,.js-referee-report").each(function(index, el) { var reportId = $(el).attr("data-reportid"), reportCount = reportIds[reportId] || 0; $(el).find(".comments-count-container,.js-referee-report-views").html(reportCount); }); var uuidInput = $("#article_uuid"), oldUUId = uuidInput.val(), newUUId = "3eb888aa-8846-41a9-b564-c4ca48b6a0db"; uuidInput.val(newUUId); $("a[href*='article_uuid=']").each(function(index, el) { var newHref = $(el).attr("href").replace(oldUUId, newUUId); $(el).attr("href", newHref); }); }); An innovative open access publishing platform offering rapid publication and open peer review, whilst supporting data deposition and sharing. Browse Gateways Collections How it Works Contact For Developers Cookie Notice Privacy Notice RSS Submit Your Research Follow us © 2012-2026 F1000 Research Ltd. ISSN 2046-1402 | Legal | Partner of Research4Life • CrossRef • ORCID • FAIRSharing R.templateTests.simpleTemplate = R.template(' $text $text $text $text $text '); R.templateTests.runTests(); var F1000platform = new F1000.Platform({ name: "f1000research", displayName: "F1000Research", hostName: "f1000research.com", id: "1", editorialEmail: "[email protected]", infoEmail: "[email protected]", usePmcStats: true }); $(function(){R.ui.dropdowns('.dropdown-for-authors, .dropdown-for-about, .dropdown-for-myresearch');}); // $(function(){R.ui.dropdowns('.dropdown-for-referees');}); $(document).ready(function () { if ($(".cookie-warning").is(":visible")) { $(".sticky").css("margin-bottom", "35px"); $(".devices").addClass("devices-and-cookie-warning"); } $(".cookie-warning .close-button").click(function (e) { $(".devices").removeClass("devices-and-cookie-warning"); $(".sticky").css("margin-bottom", "0"); }); $("#tweeter-feed .tweet-message").each(function (i, message) { var self = $(message); self.html(linkify(self.html())); }); $(".partner").on("mouseenter mouseleave", function() { $(this).find(".gray-scale, .colour").toggleClass("is-hidden"); }); }); Sign In Remember me Forgotten your password? Sign In Cancel Email or password not correct. Please try again Please wait... $(function(){ // Note: All the setup needs to run against a name attribute and *not* the id due the clonish // nature of facebox... $("a[id=googleSignInButton]").click(function(event){ event.preventDefault(); $("input[id=oAuthSystem]").val("GOOGLE"); $("form[id=oAuthForm]").submit(); }); $("a[id=facebookSignInButton]").click(function(event){ event.preventDefault(); $("input[id=oAuthSystem]").val("FACEBOOK"); $("form[id=oAuthForm]").submit(); }); $("a[id=orcidSignInButton]").click(function(event){ event.preventDefault(); $("input[id=oAuthSystem]").val("ORCID"); $("form[id=oAuthForm]").submit(); }); }); If you've forgotten your password, please enter your email address below and we'll send you instructions on how to reset your password. The email address should be the one you originally registered with F1000. Email address not valid, please try again You registered with F1000 via Google, so we cannot reset your password. To sign in, please click here . If you still need help with your Google account password, please click here . You registered with F1000 via Facebook, so we cannot reset your password. To sign in, please click here . If you still need help with your Facebook account password, please click here . Code not correct, please try again Reset password Cancel Email us for further assistance. Server error, please try again. If your email address is registered with us, we will email you instructions to reset your password. If you think you should have received this email but it has not arrived, please check your spam filters and/or contact for further assistance. Please wait... Register $(document).ready(function () { signIn.createSignInAsRow($("#sign-in-form-gfb-popup")); $(".target-field").each(function () { var uris = $(this).val().split("/"); if (uris.pop() === "login") { $(this).val(uris.toString().replace(",","/")); } }); });

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00