IDP clustering with Terracotta 3.6

Russell Beall beall at usc.edu
Thu Dec 22 22:58:50 GMT 2011


Sorry for the delay on this.

I've written the instructions for improving the EventingMapBasedStorageService to use the ConcurrentDistributedMap object.  It got to be long enough that I pulled it out of this e-mail that I was writing and added it to the end of the IdPClusterIssues page:

<https://wiki.shibboleth.net/confluence/display/SHIB2/IdPClusterIssues#IdPClusterIssues-ConfiguringtheTerracottaConcurrentDistributedMap>

For those still using TC, the performance improvement is definitely worth considering.

Regards,
Russ.

On Dec 15, 2011, at 11:59 AM, Christopher Bongaarts wrote:

> On 12/15/2011 1:18 PM, Russell Beall wrote:
> 
>> I alleviated this and other issues by converting the
>> EventingMapBasedStorageService to use the Terracotta recommended
>> ConcurrentDistributedHashMap instead of the java built-in
>> ConcurrentHashMap.  It is a simple tweak to the IdP code actually and
> [...]
>> I can publish the patch for this if desired.
> 
> Yes please (or send it to me privately), assuming it is more involved 
> than simply changing the inheritance of the EMBSS class.
> 
>> I currently have been staging and load testing a RedHat Linux cluster
>> of 2.3.3 IdP nodes clustered with 3.5.1 TC.
> 
> We have been doing load testing with RedHat Linux with 2.3.5 against TC 
> 3.5.3 (current production is IdP 2.1.3, TC 3.5.2).
> 
>> I encountered some problems with the settings currently published and
>> had to tweak them to get it to stabilize.  In particular the setting
>> which specifies the pause interval for the garbage collector became a
>> real problem and I had to get rid of it: -XX:MaxGCPauseMillis=5000
>> 
>> On Linux, this setting started being used with a vengeance (whereas
>> it basically was ignored in Solaris) and my TC server would enter
>> continuous Full GC cycles even though the heap was not full, trying
>> to make sure that it always did its GC at 5000 ms or less.  This
>> resulted in severe unworkability because really it ended up keeping
>> itself in one continuous stream of GC.
> 
> We saw something like this when I added that option to the Terracotta 
> JVM on Solaris in production.  After restarting Tomcat, the load on the 
> machine would spike up to around 5-6 (compared to a normal load of 1-2). 
>  But the system was still responding acceptably, so I let it run. 
> After a day or two, the load had settled back down to 1-2.  My current 
> conclusion is that it takes a while for the JVM to figure out the sweet 
> spot, and if it gets behind it stays behind until the load eases up 
> enough to catch up.  Also, it's an older rev of the JVM (Sun 1.6.0_13) 
> so it might be already-fixed bugs.
> 
> What are your current JVM options set to in the Linux world?  Our 
> Solaris Terracotta servers are set to:
> 
> -Xmx2g
> -XX:+UseParallelOldGC
> -XX:MaxGCPauseMillis=5000
> -verbose:gc
> -XX:+PrintGCDetails
> -XX:+PrintGCDateStamps
> -XX:-TraceClassUnloading
> (plus 64-bit options added by init script)
> 
> and Tomcat:
> 
> -d64
> -server
> -XX:+DisableExplicitGC
> -XX:+UseParallelOldGC
> -Xmx2g
> -XX:MaxPermSize=256m
> -XX:+HeapDumpOnOutOfMemoryError
> (plus terracotta boot jar/JMX options)
> 
> -- 
> %%  Christopher A. Bongaarts   %%  cab at umn.edu          %%
> %%  OIT - Identity Management  %%  http://umn.edu/~cab  %%
> %%  University of Minnesota    %%  +1 (612) 625-1809    %%
> --
> To unsubscribe from this list send an email to users-unsubscribe at shibboleth.net



More information about the users mailing list